A modern maximum-likelihood theory for high-dimensional logistic regression
From MaRDI portal
Publication:5218552
Abstract: Every student in statistics or data science learns early on that when the sample size largely exceeds the number of variables, fitting a logistic model produces estimates that are approximately unbiased. Every student also learns that there are formulas to predict the variability of these estimates which are used for the purpose of statistical inference; for instance, to produce p-values for testing the significance of regression coefficients. Although these formulas come from large sample asymptotics, we are often told that we are on reasonably safe grounds when is large in such a way that or . This paper shows that this is far from the case, and consequently, inferences routinely produced by common software packages are often unreliable. Consider a logistic model with independent features in which and become increasingly large in a fixed ratio. Then we show that (1) the MLE is biased, (2) the variability of the MLE is far greater than classically predicted, and (3) the commonly used likelihood-ratio test (LRT) is not distributed as a chi-square. The bias of the MLE is extremely problematic as it yields completely wrong predictions for the probability of a case based on observed values of the covariates. We develop a new theory, which asymptotically predicts (1) the bias of the MLE, (2) the variability of the MLE, and (3) the distribution of the LRT. We empirically also demonstrate that these predictions are extremely accurate in finite samples. Further, an appealing feature is that these novel predictions depend on the unknown sequence of regression coefficients only through a single scalar, the overall strength of the signal. This suggests very concrete procedures to adjust inference; we describe one such procedure learning a single parameter from data and producing accurate inference
Recommendations
- Maximum likelihood estimation in logistic regression models with a diverging number of covariates
- The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled Chi-square
- The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression
- Accuracies in the theory of logistic models
- Rate of convergence of the probability of non-existence of the MLE's in simple logistic regression
Cited in
(95)- Consistency of logistic regression coefficient estimates calculated from a training sample.
- Maximum likelihood estimation in logistic regression models with a diverging number of covariates
- Mallows criterion for heteroskedastic linear regressions with many regressors
- Multicarving for high-dimensional post-selection inference
- The distribution of the Lasso: uniform control over sparse balls and adaptive parameter tuning
- Some perspectives on inference in high dimensions
- Optimal combination of linear and spectral estimators for generalized linear models
- Precise statistical analysis of classification accuracies for adversarial training
- Directional testing for high dimensional multivariate normal distributions
- Fundamental barriers to high-dimensional regression with convex penalties
- Approximate message passing algorithms for rotationally invariant matrices
- The asymptotic distribution of the MLE in high-dimensional logistic models: arbitrary covariance
- Probabilistic learning inference of boundary value problem with uncertainties based on Kullback-Leibler divergence under implicit constraints
- A precise high-dimensional asymptotic theory for boosting and minimum-\(\ell_1\)-norm interpolated classifiers
- The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression
- Hierarchical inference for genome-wide association studies: a view on methodology with software
- The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models
- Which bridge estimator is the best for variable selection?
- Finite-sample analysis of \(M\)-estimators using self-concordance
- The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled Chi-square
- A regularization-based adaptive test for high-dimensional GLMs
- Online stochastic gradient descent on non-convex losses from high-dimensional inference
- scientific article; zbMATH DE number 7370646 (Why is no real title available?)
- Global and Simultaneous Hypothesis Testing for High-Dimensional Logistic Regression Models
- scientific article; zbMATH DE number 7626769 (Why is no real title available?)
- scientific article; zbMATH DE number 7625184 (Why is no real title available?)
- Approximate message passing with spectral initialization for generalized linear models*
- Analysis of overfitting in the regularized Cox model
- Binary classification of Gaussian mixtures: abundance of support vectors, benign overfitting, and regularization
- Learning from binary multiway data: probabilistic tensor decomposition and its statistical optimality
- Consistency of semi-supervised learning algorithms on graphs: probit and one-hot methods
- A Unifying Tutorial on Approximate Message Passing
- Replica analysis of overfitting in generalized linear regression models
- Penalization-induced shrinking without rotation in high dimensional GLM regression: a cavity analysis
- Sharp global convergence guarantees for iterative nonconvex optimization with random data
- Automatic bias correction for testing in high‐dimensional linear models
- Assessing the Most Vulnerable Subgroup to Type II Diabetes Associated with Statin Usage: Evidence from Electronic Health Record Data
- A Scale-Free Approach for False Discovery Rate Control in Generalized Linear Models
- Comment on “A Scale-Free Approach for False Discovery Rate Control in Generalized Linear Models” by Chengguang Dai, Buyu Lin, Xin Xing, and Jun S. Liu
- Discussion on: “A Scale-Free Approach for False Discovery Rate Control in Generalized Linear Models” by Dai, Lin, Zing, Liu
- Comments on “A Scale-Free Approach for False Discovery Rate Control in Generalized Linear Models”
- Debiased lasso for generalized linear models with a diverging number of covariates
- Knockoffs with side information
- Statistical Inference for High-Dimensional Generalized Linear Models With Binary Outcomes
- Debiasing convex regularized estimators and interval estimation in linear models
- A Friendly Tutorial on Mean-Field Spin Glass Techniques for Non-Physicists
- Universality of approximate message passing with semirandom matrices
- Online inference in high-dimensional generalized linear models with streaming data
- Dimension-agnostic inference using cross U-statistics
- Conjugate priors and bias reduction for logistic regression models
- Noisy linear inverse problems under convex constraints: exact risk asymptotics in high dimensions
- Universality of regularized regression estimators in high dimensions
- The Lasso with general Gaussian designs with applications to hypothesis testing
- StarTrek: combinatorial variable selection with false discovery rate control
- Tractability from overparametrization: the example of the negative perceptron
- Exact convergence analysis for metropolis–hastings independence samplers in Wasserstein distances
- A tradeoff between false discovery and true positive proportions for sparse high-dimensional logistic regression
- A discussion of ``A note on universal inference by Tse and Davison
- FDR control and power analysis for high-dimensional logistic regression via Stabkoff
- Bounded-memory adjusted scores estimation in generalized linear models with large data sets
- A comprehensive review of bias reduction methods for logistic regression
- Approximate message passing with rigorous guarantees for pooled data and quantitative group testing
- On the functional regression model and its finite-dimensional approximations
- An adaptively resized parametric bootstrap for inference in high-dimensional generalized linear models
- Kronecker-product random matrices and a matrix least squares problem
- Asymptotic Behavior of Adversarial Training Estimator under ℓ ∞ -Perturbation
- Precise high-dimensional asymptotics for quantifying heterogeneous transfers
- Estimating high dimensional monotone index models by iterative convex optimization
- Causal machine learning methods and use of cross-fitting in settings with high-dimensional confounding
- Efficient data integration under prior probability shift
- Improved dimension dependence in the Bernstein-von Mises theorem via a new Laplace approximation bound
- Performance of Bayesian linear regression in a model with mismatch
- Spectral estimators for structured generalized linear models via approximate message passing
- Selective inference using randomized group Lasso estimators for general models
- Entrywise dynamics and universality of general first order methods
- A leave-one-out approach to approximate message passing
- Asymptotic behaviour of the modified likelihood root
- Belief in dependence: leveraging atomic linearity in data bits for rethinking generalized linear models
- Testing sufficiency for transfer learning
- The generalization error of max-margin linear classifiers: benign overfitting and high dimensional asymptotics in the overparametrized regime
- A new central limit theorem for the augmented IPW estimator: variance inflation, cross-fit covariance and beyond
- Observable adjustments in single-index models for regularized M-estimators with bounded p/n
- Error estimation and adaptive tuning for unregularized robust M-estimator
- High dimensional logistic regression under network dependence
- A new p-value based multiple testing procedure for generalized linear models
- Debiased Lasso after sample splitting for estimation and inference in high-dimensional generalized linear models
- Equivalence of state equations from different methods in high-dimensional regression
- Spectrum-aware debiasing: a modern inference framework with applications to principal components regression
- Precise asymptotics of bagging regularized M-estimators
- High-dimensional robust regression under heavy-tailed data: asymptotics and universality
- Analysis of high-dimensional Gaussian labeled-unlabeled mixture model via message-passing algorithm
- Differentially private learning beyond the classical dimensionality regime
- Dimension-free uniform concentration bound for logistic regression
- Universality of estimators for high-dimensional linear models with block dependency
- Diaconis-Ylvisaker prior penalized likelihood for p/n(0, 1) logistic regression
This page was built for publication: A modern maximum-likelihood theory for high-dimensional logistic regression
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5218552)