Sure independence screening for ultrahigh dimensional feature space. With discussion and authors' reply
From MaRDI portal
(Redirected from Publication:4632602)
Abstract: Variable selection plays an important role in high dimensional statistical modeling which nowadays appears in many areas and is key to various scientific discoveries. For problems of large scale or dimensionality , estimation accuracy and computational cost are two top concerns. In a recent paper, Candes and Tao (2007) propose the Dantzig selector using regularization and show that it achieves the ideal risk up to a logarithmic factor . Their innovative procedure and remarkable result are challenged when the dimensionality is ultra high as the factor can be large and their uniform uncertainty principle can fail. Motivated by these concerns, we introduce the concept of sure screening and propose a sure screening method based on a correlation learning, called the Sure Independence Screening (SIS), to reduce dimensionality from high to a moderate scale that is below sample size. In a fairly general asymptotic framework, the correlation learning is shown to have the sure screening property for even exponentially growing dimensionality. As a methodological extension, an iterative SIS (ISIS) is also proposed to enhance its finite sample performance. With dimension reduced accurately from high to below sample size, variable selection can be improved on both speed and accuracy, and can then be accomplished by a well-developed method such as the SCAD, Dantzig selector, Lasso, or adaptive Lasso. The connections of these penalized least-squares methods are also elucidated.
Recommendations
- Sure independence screening in generalized linear models with NP-dimensionality
- Ultrahigh dimensional feature selection: beyond the linear model
- Factor profiled sure independence screening
- High-dimensional variable selection
- The Dantzig selector: statistical estimation when \(p\) is much larger than \(n\). (With discussions and rejoinder).
Cites work
- ``Preconditioning for feature selection and regression in high-dimensional problems
- A decision-theoretic generalization of on-line learning and an application to boosting
- A limit theorem for the norm of random matrices
- A Statistical View of Some Chemometrics Regression Tools
- Approximation and learning by greedy algorithms
- Asymptotic properties of bridge estimators in sparse high-dimensional regression models
- Asymptotics for Lasso-type estimators.
- Best subset selection, persistence in high-dimensional statistical learning and optimization under l₁ constraint
- Better Subset Regression Using the Nonnegative Garrote
- Comments on: ``Wavelets in statistics: a review by A. Antoniadis
- Deviation Inequalities on Largest Eigenvalues
- Geometric Representation of High Dimension, Low Sample Size Data
- Heuristics of instability and stabilization in model selection
- High-dimensional classification using features annealed independence rules
- High-dimensional graphs and variable selection with the Lasso
- scientific article; zbMATH DE number 5957408 (Why is no real title available?)
- scientific article; zbMATH DE number 3988038 (Why is no real title available?)
- scientific article; zbMATH DE number 51418 (Why is no real title available?)
- scientific article; zbMATH DE number 1347881 (Why is no real title available?)
- scientific article; zbMATH DE number 1034042 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- Ideal spatial adaptation by wavelet shrinkage
- Least angle regression. (With discussion)
- Limit of the smallest eigenvalue of a large dimensional sample covariance matrix
- Local Strong Homogeneity of a Regularized Estimator
- Nonconcave penalized likelihood with a diverging number of parameters.
- On the distribution of the largest eigenvalue in principal components analysis
- Optimally sparse representation in general (nonorthogonal) dictionaries via ℓ 1 minimization
- Pathwise coordinate optimization
- Persistene in high-dimensional linear predictor-selection and the virtue of overparametrization
- Regularization of Wavelet Approximations
- Regularized estimation of large covariance matrices
- Rejoinder: One-step sparse estimates in nonconcave penalized likelihood models
- Relaxed Lasso
- Simultaneous analysis of Lasso and Dantzig selector
- Some theory for Fisher's linear discriminant function, `naive Bayes', and some alternatives when there are many more variables than observations
- Sparsistency and rates of convergence in large covariance matrix estimation
- Statistical challenges with high dimensionality: feature selection in knowledge discovery
- Statistical significance for genomewide studies
- Statistics on special manifolds
- The Adaptive Lasso and Its Oracle Properties
- The concentration of measure phenomenon
- The Dantzig selector: statistical estimation when \(p\) is much larger than \(n\). (With discussions and rejoinder).
- The Group Lasso for Logistic Regression
- The smallest eigenvalue of a large dimensional Wishart matrix
- The sparsity and bias of the LASSO selection in high-dimensional linear regression
- Uncertainty principles and ideal atomic decomposition
- Variable selection for Cox's proportional hazards model and frailty model
- Variable selection using MM algorithms
- Variable Selection via Nonconcave Penalized Likelihood and its Oracle Properties
- Weak convergence and empirical processes. With applications to statistics
Cited in
(only showing first 100 items - show all)- Independent rule in classification of multivariate binary data
- Stabilizing variable selection and regression
- Selection by partitioning the solution paths
- Nearly unbiased variable selection under minimax concave penalty
- Sparse classification with paired covariates
- Factor-Adjusted Regularized Model Selection
- Restricted fence method for covariate selection in longitudinal data analysis
- Nonparametric feature screening
- On selecting interacting features from high-dimensional data
- Using random subspace method for prediction and variable importance assessment in linear regression
- LOL selection in high dimension
- Group subset selection for linear regression
- A sequential test for variable selection in high dimensional complex data
- Network-based feature screening with applications to genome data
- Individual-level social influence identification in social media: a learning-simulation coordinated method
- Variable selection in censored quantile regression with high dimensional data
- Robust dependence measure for detecting associations in large data set
- Conditional feature screening for mean and variance functions in models with multiple-index structure
- Censored cumulative residual independent screening for ultrahigh-dimensional survival data
- Nonparametric independence screening via favored smoothing bandwidth
- Conditional quantile correlation screening procedure for ultrahigh-dimensional varying coefficient models
- Powerful test based on conditional effects for genome-wide screening
- Test for high-dimensional regression coefficients using refitted cross-validation variance estimation
- Are discoveries spurious? Distributions of maximum spurious correlations and their applications
- Screening group variables in the proportional hazards model
- Ultrahigh dimensional feature screening via projection
- Two-layer EM algorithm for ALD mixture regression models: a new solution to composite quantile regression
- Principal components adjusted variable screening
- Correlation rank screening for ultrahigh-dimensional survival data
- Variable selection using shrinkage priors
- Canonical kernel dimension reduction
- Model free feature screening for ultrahigh dimensional data with responses missing at random
- Feature screening for generalized varying coefficient models with application to dichotomous responses
- Adaptive conditional feature screening
- Jackknife empirical likelihood test for high-dimensional regression coefficients
- High-dimensional multivariate posterior consistency under global-local shrinkage priors
- A new nonparametric screening method for ultrahigh-dimensional survival data
- Robust feature screening for ultra-high dimensional right censored data via distance correlation
- Fused mean-variance filter for feature screening
- Variable selection for high dimensional Gaussian copula regression model: an adaptive hypothesis testing procedure
- Sparse pathway-based prediction models for high-throughput molecular data
- On the oracle property of a generalized adaptive elastic-net for multivariate linear regression with a diverging number of parameters
- Adjusted Pearson chi-square feature screening for multi-classification with ultrahigh dimensional data
- Conditional screening for ultra-high dimensional covariates with survival outcomes
- Model-free conditional independence feature screening for ultrahigh dimensional data
- Model-free feature screening for ultrahigh dimensional censored regression
- Extended differential geometric LARS for high-dimensional GLMs with general dispersion parameter
- Asymptotically honest confidence regions for high dimensional parameters by the desparsified conservative Lasso
- Optimal directional statistic for general regression
- Robust conditional nonparametric independence screening for ultrahigh-dimensional data
- A simple model-free survival conditional feature screening
- Model selection using mass-nonlocal prior
- Covariance-insured screening
- Feature screening in ultrahigh-dimensional partially linear models with missing responses at random
- Regression adjustment for treatment effect with multicollinearity in high dimensions
- Modified SCAD penalty for constrained variable selection problems
- Nonconvex penalized ridge estimations for partially linear additive models in ultrahigh dimension
- Feature elimination in kernel machines in moderately high dimensions
- An RKHS model for variable selection in functional linear regression
- Determination of vector error correction models in high dimensions
- Portal nodes screening for large scale social networks
- Stable feature screening for ultrahigh dimensional data
- Variable screening for ultrahigh dimensional heterogeneous data via conditional quantile correlations
- Model-free feature screening for ultrahigh-dimensional data conditional on some variables
- Variable screening for high dimensional time series
- Hypothesis testing sure independence screening for nonparametric regression
- Conditional mean and quantile dependence testing in high dimension
- Principles of experimental design for big data analysis
- On consistency and sparsity for sliced inverse regression in high dimensions
- I-LAMM for sparse learning: simultaneous control of algorithmic complexity and statistical error
- A retail store SKU promotions optimization model for category multi-period profit maximization
- Optimal estimation of slope vector in high-dimensional linear transformation models
- Feature screening for nonparametric and semiparametric models with ultrahigh-dimensional covariates
- Nonparametric independence feature screening for ultrahigh-dimensional survival data
- High-dimensional inference: confidence intervals, \(p\)-values and R-software \texttt{hdi}
- An RKHS-based approach to double-penalized regression in high-dimensional partially linear models
- Broken adaptive ridge regression and its asymptotic properties
- Factor-adjusted multiple testing of correlations
- Feature screening for multi-response varying coefficient models with ultrahigh dimensional predictors
- Optimal feature selection for sparse linear discriminant analysis and its applications in gene expression data
- Change-point detection in multinomial data with a large number of categories
- An iterative algorithm for fitting nonconvex penalized generalized linear models with grouped predictors
- Model-free feature screening for high-dimensional survival data
- Conditional-quantile screening for ultrahigh-dimensional survival data via martingale difference correlation
- Debiasing the Lasso: optimal sample size for Gaussian designs
- Measuring and testing for interval quantile dependence
- Beyond Gaussian approximation: bootstrap for maxima of sums of independent random vectors
- A method for selecting the relevant dimensions for high-dimensional classification in singular vector spaces
- Fused variable screening for massive imbalanced data
- Feature screening for ultrahigh dimensional categorical data with covariates missing at random
- A nonparametric feature screening method for ultrahigh-dimensional missing response
- Approximate least squares estimation for spatial autoregressive models with covariates
- A note on quantile feature screening via distance correlation
- Low-dimensional confounder adjustment and high-dimensional penalized estimation for survival analysis
- Determining cutoff point of ensemble trees based on sample size in predicting clinical dose with DNA microarray data
- A distribution-based Lasso for a general single-index model
- Variable selection for partially linear models via Bayesian subset modeling with diffusing prior
- Penalized empirical likelihood for partially linear errors-in-variables panel data models with fixed effects
- Feature screening based on distance correlation for ultrahigh-dimensional censored data with covariate measurement error
- An efficient algorithm for joint feature screening in ultrahigh-dimensional Cox's model
This page was built for publication: Sure independence screening for ultrahigh dimensional feature space. With discussion and authors' reply
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4632602)