High-dimensional variable selection
From MaRDI portal
Point estimation (62F10) Asymptotic properties of parametric estimators (62F12) Estimation in multivariate analysis (62H12) Linear regression; mixed models (62J05) Ridge regression; shrinkage estimators (Lasso) (62J07) Applications of statistics to biology and medical sciences; meta analysis (62P10)
Abstract: This paper explores the following question: what kind of statistical guarantees can be given when doing variable selection in high-dimensional models? In particular, we look at the error rates and power of some multi-stage regression methods. In the first stage we fit a set of candidate models. In the second stage we select one model by cross-validation. In the third stage we use hypothesis testing to eliminate some variables. We refer to the first two stages as "screening" and the last stage as "cleaning." We consider three screening methods: the lasso, marginal regression, and forward stepwise regression. Our method gives consistent variable selection under certain conditions.
Recommendations
- Variable selection in high dimensional data analysis with applications
- Variable selection for high dimensional multivariate outcomes
- Variable selection in high-dimensional partially linear models
- A Selective Overview of Variable Selection in High Dimensional Feature Space (Invited Review Article)
- Selection of variables and dimension reduction in high-dimensional non-parametric regression
- Variable selection methods in high-dimensional regression -- a simulation study
- Simultaneous dimension reduction and variable selection in modeling high dimensional data
- Variable selection and estimation in high-dimensional partially linear models
- High Dimensional Variable Selection via Tilting
- Feature selection for high-dimensional data
Cites work
- Approximation and learning by greedy algorithms
- Boosting for high-dimensional linear models
- Causation, prediction, and search. With additional material by David Heckerman, Christopher Meek, Gregory F. Cooper and Thomas Richardson.
- For most large underdetermined systems of linear equations the minimal 𝓁1‐norm solution is also the sparsest solution
- Greed is Good: Algorithmic Results for Sparse Approximation
- High-dimensional graphs and variable selection with the Lasso
- scientific article; zbMATH DE number 5957408 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- Just relax: convex programming methods for identifying sparse signals in noise
- Lasso-type recovery of sparse representations for high-dimensional data
- Least angle regression. (With discussion)
- Persistene in high-dimensional linear predictor-selection and the virtue of overparametrization
- Relaxed Lasso
- The Adaptive Lasso and Its Oracle Properties
- The Dantzig selector: statistical estimation when \(p\) is much larger than \(n\). (With discussions and rejoinder).
- The sparsity and bias of the LASSO selection in high-dimensional linear regression
- Uniform consistency in causal inference
Cited in
(only showing first 100 items - show all)- Confidence intervals for high-dimensional inverse covariance estimation
- High-dimensional regression and variable selection using CAR scores
- False Discovery Rate Control Under General Dependence By Symmetrized Data Aggregation
- LOL selection in high dimension
- Asymptotics for high dimensional regression \(M\)-estimates: fixed design results
- A unified theory of confidence regions and testing for high-dimensional estimating equations
- Simultaneous dimension reduction and variable selection in modeling high dimensional data
- Robust stability best subset selection for autocorrelated data based on robust location and dispersion estimator
- Principal components adjusted variable screening
- High-dimensional simultaneous inference with the bootstrap
- Convex and non-convex regularization methods for spatial point processes intensity estimation
- Efficient test-based variable selection for high-dimensional linear models
- Selective inference with a randomized response
- Projection tests for high-dimensional spiked covariance matrices
- High-dimensional inference: confidence intervals, \(p\)-values and R-software \texttt{hdi}
- Factor-adjusted multiple testing of correlations
- Variable selection with Hamming loss
- Honest variable selection in linear and logistic regression models via \(\ell _{1}\) and \(\ell _{1}+\ell _{2}\) penalization
- The adaptive and the thresholded Lasso for potentially misspecified models (and a lower bound for the Lasso)
- Debiasing the Lasso: optimal sample size for Gaussian designs
- Inference under Fine-Gray competing risks model with high-dimensional covariates
- Multicarving for high-dimensional post-selection inference
- Exact model comparisons in the plausibility framework
- Sure independence screening in the presence of missing data
- In defense of the indefensible: a very naïve approach to high-dimensional inference
- Mining events with declassified diplomatic documents
- Spatially relaxed inference on high-dimensional linear models
- Network differential connectivity analysis
- Projection-based high-dimensional sign test
- Iterative algorithm for discrete structure recovery
- Self-semi-supervised clustering for large scale data with massive null group
- Post-model-selection inference in linear regression models: an integrated review
- Thresholding tests based on affine Lasso to achieve non-asymptotic nominal level and high power under sparse and dense alternatives in high dimension
- Hierarchical inference for genome-wide association studies: a view on methodology with software
- Debiasing the debiased Lasso with bootstrap
- Fundamental limits of exact support recovery in high dimensions
- Which bridge estimator is the best for variable selection?
- Exact tests via multiple data splitting
- Variable selection techniques after multiple imputation in high-dimensional data
- Dynamic tilted current correlation for high dimensional variable screening
- A significance test for the lasso
- Discussion: ``A significance test for the lasso
- Rejoinder: ``A significance test for the lasso
- High-dimensional variable screening and bias in subsequent inference, with an empirical comparison
- A global homogeneity test for high-dimensional linear regression
- Feature selection for high-dimensional data
- Bootstrapping and sample splitting for high-dimensional, assumption-lean inference
- Inference for L₂-boosting
- Selective inference via marginal screening for high dimensional classification
- A scalable nonparametric specification testing for massive data
- Spectral analysis of high-dimensional time series
- A knockoff filter for high-dimensional selective inference
- On the impact of model selection on predictor identification and parameter inference
- Optimal two-step prediction in regression
- Predictor ranking and false discovery proportion control in high-dimensional regression
- Variable selection procedures from multiple testing
- Empirical likelihood test for high dimensional linear models
- Variable screening in predicting clinical outcome with high-dimensional microarrays
- Endogeneity in high dimensions
- Tolerance intervals from ridge regression in the presence of multicollinearity and high dimension
- Tests for high-dimensional single-index models
- Detection of gene-gene interactions using multistage sparse and low-rank regression
- Variable selection using stepdown procedures in high-dimensional linear models
- Estimation for high-dimensional linear mixed-effects models using _1-penalization
- Optimality of Graphlet Screening in High Dimensional Variable Selection
- Statistical learning and selective inference
- The benefit of group sparsity in group inference with de-biased scaled group Lasso
- Screening-based Bregman divergence estimation with NP-dimensionality
- Thresholding least-squares inference in high-dimensional regression models
- Two-stage procedures for high-dimensional data
- Integrative analysis and variable selection with multiple high-dimensional data sets
- The predictive power of the business and bank sentiment of firms: a high-dimensional Granger causality approach
- Feature selection in finite mixture of sparse normal linear models in high-dimensional feature space
- Debiased Inference on Treatment Effect in a High-Dimensional Model
- Selection of the Regularization Parameter in Graphical Models Using Network Characteristics
- A Selective Overview of Variable Selection in High Dimensional Feature Space (Invited Review Article)
- Statistical significance in high-dimensional linear models
- The geometry of least squares in the 21st century
- Classifier variability: accounting for training and testing
- Sharp support recovery from noisy random measurements by \(\ell_1\)-minimization
- scientific article; zbMATH DE number 736274 (Why is no real title available?)
- UPS delivers optimal phase diagram in high-dimensional variable selection
- Estimation and Inference of Heterogeneous Treatment Effects using Random Forests
- Goodness-of-Fit Tests for High Dimensional Linear Models
- Sure independence screening for ultrahigh dimensional feature space. With discussion and authors' reply
- High Dimensional Variable Selection via Tilting
- Bayesian high-dimensional screening via MCMC
- Covariate assisted screening and estimation
- Penalized weighted composite quantile regression in the linear regression model with heavy-tailed autocorrelated errors
- Selecting massive variables using an iterated conditional modes/medians algorithm
- High-dimensional inference in misspecified linear models
- Two-sample spatial rank test using projection
- Variable selection for longitudinal data with high-dimensional covariates and dropouts
- A regularization-based adaptive test for high-dimensional GLMs
- Analysis of testing-based forward model selection
- Hypothesis testing in large-scale functional linear regression
- The revisited knockoffs method for variable selection in L1-penalized regressions
- Variable Selection With Second-Generation P-Values
- scientific article; zbMATH DE number 7626707 (Why is no real title available?)
- An ensemble learning method for variable selection: application to high-dimensional data and missing values
This page was built for publication: High-dimensional variable selection
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q834336)