Nonpenalized variable selection in high-dimensional linear model settings via generalized fiducial inference

From MaRDI portal



Abstract: Standard penalized methods of variable selection and parameter estimation rely on the magnitude of coefficient estimates to decide which variables to include in the final model. However, coefficient estimates are unreliable when the design matrix is collinear. To overcome this challenge an entirely new perspective on variable selection is presented within a generalized fiducial inference framework. This new procedure is able to effectively account for linear dependencies among subsets of covariates in a high-dimensional setting where p can grow almost exponentially in n, as well as in the classical setting where plen. It is shown that the procedure very naturally assigns small probabilities to subsets of covariates which include redundancies by way of explicit L0 minimization. Furthermore, with a typical sparsity assumption, it is shown that the proposed method is consistent in the sense that the probability of the true sparse subset of covariates converges in probability to 1 as noinfty, or as noinfty and poinfty. Very reasonable conditions are needed, and little restriction is placed on the class of possible subsets of covariates to achieve this consistency result.


This article develops a new perspective on variable selection to deal with linear dependences among subsets of covariates in a high-dimensional setting. A generalized fiducial inference framework is adopted. A procedure to assign small probabilities to subsets of covariates which include redundancies by way of explicit Lo minimization using a Dantzig selector approach is presented. A concept of $\varepsilon$-admissible subsets is proposed to give rise to the determination of a posterior-like distribution which assigns negligible probabilities to all subsets with redundancies in the true data generating model. It is shown that, under a typical sparsity condition, the probability of the true data generating model converges to 1. Comparisons with Bayesian and frequentist methods are performed in two simulation setups.











This page was built for publication: Nonpenalized variable selection in high-dimensional linear model settings via generalized fiducial inference

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2414104)