Dimension reduction and variable selection in case control studies via regularized likelihood optimization
From MaRDI portal
Abstract: Dimension reduction and variable selection are performed routinely in case-control studies, but the literature on the theoretical aspects of the resulting estimates is scarce. We bring our contribution to this literature by studying estimators obtained via L1 penalized likelihood optimization. We show that the optimizers of the L1 penalized retrospective likelihood coincide with the optimizers of the L1 penalized prospective likelihood. This extends the results of Prentice and Pyke (1979), obtained for non-regularized likelihoods. We establish both the sup-norm consistency of the odds ratio, after model selection, and the consistency of subset selection of our estimators. The novelty of our theoretical results consists in the study of these properties under the case-control sampling scheme. Our results hold for selection performed over a large collection of candidate variables, with cardinality allowed to depend and be greater than the sample size. We complement our theoretical results with a novel approach of determining data driven tuning parameters, based on the bisection method. The resulting procedure offers significant computational savings when compared with grid search based methods. All our numerical experiments support strongly our theoretical findings.
Recommendations
- Penalized full likelihood approach to variable selection for Cox's regression model under nested case-control sampling
- Variable Selection via Nonconcave Penalized Likelihood and its Oracle Properties
- Likelihood-based selection and sharp parameter estimation
- scientific article; zbMATH DE number 6468204
- Honest variable selection in linear and logistic regression models via \(\ell _{1}\) and \(\ell _{1}+\ell _{2}\) penalization
Cites work
- A goodness-of-fit test for logistic regression models based on case-control data
- A note on the Lasso and related procedures in model selection
- An interior-point method for large-scale l₁-regularized logistic regression
- Asymptotic inference for semiparametric association models
- Combinatorial methods in density estimation
- Covariate selection for semiparametric hazard function regression models
- High-dimensional additive modeling
- High-dimensional generalized linear models and the lasso
- High-dimensional Ising model selection using \(\ell _{1}\)-regularized logistic regression
- Honest variable selection in linear and logistic regression models via \(\ell _{1}\) and \(\ell _{1}+\ell _{2}\) penalization
- scientific article; zbMATH DE number 5957245 (Why is no real title available?)
- scientific article; zbMATH DE number 42384 (Why is no real title available?)
- Large sample theory of empirical distributions in biased sampling models
- Lasso type classifiers with a reject option
- Logistic disease incidence models and case-control studies
- On the semi-parametric efficiency of logistic regression under case-control sampling
- Piecewise linear regularized solution paths
- Prospective Analysis of Logistic Case-Control Studies
- Risk bounds for model selection via penalization
- Semiparametric mixtures in case-control studies
- Separate sample logistic discrimination
- Shrinkage estimators for robust and efficient inference in haplotype-based case-control studies
- Some results on the estimation of logistic models based on retrospective data
Cited in
(5)- An iterative algorithm for fitting nonconvex penalized generalized linear models with grouped predictors
- Global and local two-sample tests via regression
- Likelihood-based selection and sharp parameter estimation
- Tight conditions for consistency of variable selection in the context of high dimensionality
- SPADES and mixture models
This page was built for publication: Dimension reduction and variable selection in case control studies via regularized likelihood optimization
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1952024)