Aggregation for Gaussian regression
From MaRDI portal
Abstract: This paper studies statistical aggregation procedures in the regression setting. A motivating factor is the existence of many different methods of estimation, leading to possibly competing estimators. We consider here three different types of aggregation: model selection (MS) aggregation, convex (C) aggregation and linear (L) aggregation. The objective of (MS) is to select the optimal single estimator from the list; that of (C) is to select the optimal convex combination of the given estimators; and that of (L) is to select the optimal linear combination of the given estimators. We are interested in evaluating the rates of convergence of the excess risks of the estimators obtained by these procedures. Our approach is motivated by recently published minimax results [Nemirovski, A. (2000). Topics in non-parametric statistics. Lectures on Probability Theory and Statistics (Saint-Flour, 1998). Lecture Notes in Math. 1738 85--277. Springer, Berlin; Tsybakov, A. B. (2003). Optimal rates of aggregation. Learning Theory and Kernel Machines. Lecture Notes in Artificial Intelligence 2777 303--313. Springer, Heidelberg]. There exist competing aggregation procedures achieving optimal convergence rates for each of the (MS), (C) and (L) cases separately. Since these procedures are not directly comparable with each other, we suggest an alternative solution. We prove that all three optimal rates, as well as those for the newly introduced (S) aggregation (subset selection), are nearly achieved via a single ``universal aggregation procedure. The procedure consists of mixing the initial estimators with weights obtained by penalized least squares. Two different penalties are considered: one of them is of the BIC type, the second one is a data-dependent -type penalty.
Recommendations
Cites work
- A distribution-free theory of nonparametric regression
- A new look at the statistical model identification
- Adaptive estimation with soft thresholding penalties
- Adaptive model selection using empirical complexities
- Adaptive Regression by Mixing
- Aggregated estimators and empirical complexity for least square regression
- Aggregating regression procedures to improve performance
- Atomic decomposition by basis pursuit
- Combining different procedures for adaptive regression
- Consistent covariate selection and post model selection inference in semiparametric regression.
- Estimating the dimension of a model
- Functional aggregation for nonparametric regression.
- Gaussian model selection
- scientific article; zbMATH DE number 1522808 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- scientific article; zbMATH DE number 893887 (Why is no real title available?)
- Information Theory and Mixing Least-Squares Regressions
- Introduction to nonparametric estimation
- Learning by mirror averaging
- Learning Theory and Kernel Machines
- Least angle regression. (With discussion)
- Local Rademacher complexities and oracle inequalities in risk minimization. (2004 IMS Medallion Lecture). (With discussions and rejoinder)
- Model selection for regression on a fixed design
- Model selection for regression on a random design
- Model selection in nonparametric regression
- Model selection via testing: an alternative to (penalized) maximum likelihood estimators.
- Oracle inequalities for inverse problems
- Ordered linear smoothers
- Recursive aggregation of estimators by the mirror descent algorithm with averaging
- Regularization of Wavelet Approximations
- Risk bounds for model selection via penalization
- Sequential Procedures for Aggregating Arbitrary Estimators of a Conditional Mean
- Some Comments on C P
- Stable recovery of sparse overcomplete representations in the presence of noise
- Statistical learning theory and stochastic optimization. Ecole d'Eté de Probabilitiés de Saint-Flour XXXI -- 2001.
- The boosting approach to machine learning: an overview
- The risk inflation criterion for multiple regression
- Universal approximation bounds for superpositions of a sigmoidal function
- Wavelets, approximation, and statistical applications
Cited in
(only showing first 100 items - show all)- Lasso-type recovery of sparse representations for high-dimensional data
- A universal procedure for aggregating estimators
- Mixing least-squares estimators when the variance is unknown
- Aggregation by exponential weighting, sharp PAC-Bayesian bounds and sparsity
- Estimator selection in the Gaussian setting
- Aggregating regression procedures to improve performance
- Minimax lower bounds for the simultaneous wavelet deconvolution with fractional Gaussian noise and unknown kernels
- A new approach to estimator selection
- Optimal Kullback-Leibler aggregation in mixture density estimation by maximum likelihood
- On the sensitivity of the Lasso to the number of predictor variables
- Functional aggregation for nonparametric regression.
- Aggregated estimators and empirical complexity for least square regression
- Sharp oracle inequalities for aggregation of affine estimators
- On the optimality of the empirical risk minimization procedure for the convex aggregation problem
- High-dimensional additive hazards models and the lasso
- Model selection in regression under structural constraints
- Laplace deconvolution with noisy observations
- Optimal model selection in heteroscedastic regression using piecewise polynomial functions
- On the asymptotic properties of the group lasso estimator for linear models
- Honest variable selection in linear and logistic regression models via \(\ell _{1}\) and \(\ell _{1}+\ell _{2}\) penalization
- General oracle inequalities for model selection
- On the conditions used to prove oracle results for the Lasso
- MAP model selection in Gaussian regression
- PAC-Bayesian bounds for sparse regression estimation with exponential weights
- The adaptive and the thresholded Lasso for potentially misspecified models (and a lower bound for the Lasso)
- Sparsity considerations for dependent variables
- The smooth-Lasso and other \(\ell _{1}+\ell _{2}\)-penalized methods
- Least squares after model selection in high-dimensional sparse models
- On the optimality of the aggregate with exponential weights for low temperatures
- Sign-constrained least squares estimation for high-dimensional regression
- Robust subset selection
- Exact minimax risk for linear least squares, and the lower tail of sample covariance matrices
- Aggregating estimates by convex optimization
- Aggregation of estimators and stochastic optimization
- Pivotal estimation via square-root lasso in nonparametric regression
- Non-parametric Poisson regression from independent and weakly dependent observations by model selection
- Anisotropic functional Laplace deconvolution
- Aggregation using input-output trade-off
- Structured estimation for the nonparametric Cox model
- Lasso and probabilistic inequalities for multivariate point processes
- Adaptive estimation over anisotropic functional classes via oracle approach
- Sparse high-dimensional varying coefficient model: nonasymptotic minimax study
- Simultaneous analysis of Lasso and Dantzig selector
- On minimax convergence rates under \(L^p\)-risk for the anisotropic functional deconvolution model
- Sup-norm convergence rate and sign concentration property of Lasso and Dantzig estimators
- \(\ell_1\)-penalized quantile regression in high-dimensional sparse models
- Empirical risk minimization is optimal for the convex aggregation problem
- Multichannel deconvolution with long-range dependence: a minimax study
- Model averaging by jackknife criterion in models with dependent data
- Optimal learning with \textit{Q}-aggregation
- Simultaneous adaptation to the margin and to complexity in classification
- Optimal rates of aggregation in classification under low noise assumption
- Adaptive estimation of the baseline hazard function in the Cox model by model selection, with high-dimensional covariates
- Greedy algorithms for prediction
- Combining a relaxed EM algorithm with Occam's razor for Bayesian variable selection in high-dimensional regression
- A nonlinear aggregation type classifier
- Best subset selection via a modern optimization lens
- Performance of empirical risk minimization in linear aggregation
- SLOPE is adaptive to unknown sparsity and asymptotically minimax
- Non-asymptotic oracle inequalities for the Lasso and group Lasso in high dimensional logistic model
- Aggregated wavelet estimation and its application to ultra-fast fMRI
- Deconvolution model with fractional Gaussian noise: a minimax study
- AIC for the Lasso in generalized linear models
- Estimation of matrices with row sparsity
- Sharp connections between Berry-Esseen characteristics and Edgeworth expansions for stationary processes
- Classification of longitudinal data through a semiparametric mixed-effects model based on Lasso-type estimators
- Anisotropic de-noising in functional deconvolution model with dimension-free convergence rates
- Nonparametric sequential prediction of time series
- Optimal equivariant prediction for high-dimensional linear models with arbitrary predictor covariance
- Sparse regression learning by aggregation and Langevin Monte-Carlo
- Mirror averaging with sparsity priors
- Transductive versions of the Lasso and the Dantzig selector
- Kullback-Leibler aggregation and misspecified generalized linear models
- Laplace deconvolution with dependent errors: a minimax study
- Sharp oracle inequalities for square root regularization
- Microlocal analysis of the geometric separation problem
- Structured, sparse aggregation
- scientific article; zbMATH DE number 7625186 (Why is no real title available?)
- Minimax adaptive wavelet estimator for the anisotropic functional deconvolution model with unknown kernel
- Blind deconvolution model in periodic setting with fractional Gaussian noise
- Sparse linear regression models of high dimensional covariates with non-Gaussian outliers and Berkson error-in-variable under heteroscedasticity
- Anisotropic functional deconvolution for the irregular design: A minimax study
- Estimation in nonparametric regression model with additive and multiplicative noise via Laguerre series
- Learning sparse classifiers: continuous and mixed integer optimization perspectives
- Anisotropic functional deconvolution with long-memory noise: the case of a multi-parameter fractional Wiener sheet
- Hyper-sparse optimal aggregation
- Prediction of time series by statistical learning: general losses and fast rates
- Exponential screening and optimal rates of sparse estimation
- Estimation of high-dimensional low-rank matrices
- Aggregation of regularized solutions from multiple observation models
- Quasi-likelihood and/or robust estimation in high dimensions
- High-dimensional regression with unknown variance
- A unified framework for high-dimensional analysis of M-estimators with decomposable regularizers
- Sparse estimation by exponential weighting
- Sample average approximation with heavier tails II: localization in stochastic convex optimization and persistence results for the Lasso
- Sparse recovery under matrix uncertainty
- A Critical Review of LASSO and Its Derivatives for Variable Selection Under Dependence Among Covariates
- Model aggregation for doubly divided data with large size and large dimension
- Variance function estimation in regression model via aggregation procedures
- Simple proof of the risk bound for denoising by exponential weights for asymmetric noise distributions
This page was built for publication: Aggregation for Gaussian regression
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2456016)