Optimal subsampling for large sample logistic regression
From MaRDI portal
Abstract: For massive data, the family of subsampling algorithms is popular to downsize the data volume and reduce computational burden. Existing studies focus on approximating the ordinary least squares estimate in linear regression, where statistical leverage scores are often used to define subsampling probabilities. In this paper, we propose fast subsampling algorithms to efficiently approximate the maximum likelihood estimate in logistic regression. We first establish consistency and asymptotic normality of the estimator from a general subsampling algorithm, and then derive optimal subsampling probabilities that minimize the asymptotic mean squared error of the resultant estimator. An alternative minimization criterion is also proposed to further reduce the computational cost. The optimal subsampling probabilities depend on the full data estimate, so we develop a two-step algorithm to approximate the optimal subsampling procedure. This algorithm is computationally efficient and has a significant reduction in computing time compared to the full data approach. Consistency and asymptotic normality of the estimator from a two-step algorithm are also established. Synthetic and real data sets are used to evaluate the practical performance of the proposed method.
Recommendations
Cites work
- A fast randomized algorithm for overdetermined linear least-squares regression
- A statistical perspective on algorithmic leveraging
- Applied logistic regression
- Bayesian data analysis.
- Bootstrap methods: another look at the jackknife
- CUR matrix decompositions for improved data analysis
- Fast approximation of matrix coherence and statistical leverage
- Faster least squares approximation
- scientific article; zbMATH DE number 5957445 (Why is no real title available?)
- scientific article; zbMATH DE number 3984372 (Why is no real title available?)
- scientific article; zbMATH DE number 3176492 (Why is no real title available?)
- scientific article; zbMATH DE number 3746271 (Why is no real title available?)
- Local case-control sampling: efficient subsampling in imbalanced data sets
- Optimum experimental designs, with SAS
- Sampling algorithms for l₂ regression and applications
- Sub-Gaussian random variables
Cited in
(only showing first 100 items - show all)- Randomized sketches for kernel CCA
- Optimal subsampling for large-scale quantile regression
- Multiplicative perturbation bounds for multivariate multiple linear regression in Schatten p-norms
- Surprise sampling: improving and extending the local case-control sampling
- Parallel-and-stream accelerator for computationally fast supervised learning
- Optimal subsampling for composite quantile regression in big data
- Optimal subsampling for least absolute relative error estimators with massive data
- Model-free global likelihood subsampling for massive data
- Subdata selection algorithm for linear model discrimination
- Robust active learning with binary responses
- Optimal subsample selection for massive logistic regression with distributed data
- Score-matching representative approach for big data analysis with generalized linear models
- Functional principal subspace sampling for large scale functional data analysis
- A two-stage optimal subsampling estimation for missing data problems with large-scale data
- Inversion-free subsampling Newton's method for large sample logistic regression
- Statistical inference in massive datasets by empirical likelihood
- Optimal subsampling for composite quantile regression model in massive data
- Surface temperature monitoring in liver procurement via functional variance change-point analysis
- Information-based optimal subdata selection for big data logistic regression
- Online updating method to correct for measurement error in big data streams
- Local uncertainty sampling for large-scale multiclass logistic regression
- Testing multivariate quantile by empirical likelihood
- Crawling subsampling for multivariate spatial autoregression model in large-scale networks
- Learning nonlocal constitutive models with neural networks
- A quasi-Monte Carlo data compression algorithm for machine learning
- Bayesian estimation under informative sampling with unattenuated dependence
- Divide-and-conquer information-based optimal subdata selection algorithm
- Optimal subsampling for softmax regression
- Estimating promotion effects in email marketing using a large-scale cross-classified Bayesian joint model for nested imbalanced data
- Local case-control sampling: efficient subsampling in imbalanced data sets
- Optimal subsampling algorithms for big data regressions
- Randomized Spectral Clustering in Large-Scale Stochastic Block Models
- Optimal Sampling for Generalized Linear Models Under Measurement Constraints
- LowCon: A Design-based Subsampling Approach in a Misspecified Linear Model
- Least-Square Approximation for a Distributed System
- Logistic Regression Models for Aggregated Data
- Online updating of information based model selection in the big data setting
- More efficient estimation for logistic regression with optimal subsamples
- Optimal subsampling for quantile regression in big data
- Optimal Distributed Subsampling for Maximum Quasi-Likelihood Estimators With Massive Data
- Model Checking in Large-Scale Dataset via Structure-Adaptive-Sampling
- Optimal subsampling for large‐sample quantile regression with massive data
- Fast Calibration for Computer Models with Massive Physical Observations
- Optimal subsampling for multiplicative regression with massive data
- Subsampling spectral clustering for stochastic block models in large-scale networks
- Information-based optimal subdata selection for non-linear models
- A model robust subsampling approach for generalised linear models in big data settings
- Subsampling and Jackknifing: A Practically Convenient Solution for Large Data Analysis With Limited Computational Resources
- Sketched approximation of regularized canonical correlation analysis
- Optimal sampling algorithms for block matrix multiplication
- Model constraints independent optimal subsampling probabilities for softmax regression
- Applications of robust methods in spatial analysis
- Subdata selection based on orthogonal array for big data
- Generalized linear models for massive data via doubly-sketching
- Three-way sampling for rapid attribute reduction
- Distributed smoothed rank regression with heterogeneous errors for massive data
- A block-randomized stochastic method with importance sampling for CP tensor decomposition
- Optimal subsampling algorithms for composite quantile regression in massive data
- Optimal sampling designs for multidimensional streaming time series with application to power grid sensor data
- Optimal decorrelated score subsampling for generalized linear models with massive data
- Communication-efficient surrogate quantile regression for non-randomly distributed system
- Conditional characteristic feature screening for massive imbalanced data
- LIC criterion for optimal subset selection in distributed interval estimation
- Subsampling in longitudinal models
- Frugal Gaussian clustering of huge imbalanced datasets through a bin-marginal approach
- Analyzing Big EHR Data—Optimal Cox Regression Subsampling Procedure with Rare Events
- Optimal subsampling for functional quantile regression
- Unweighted estimation based on optimal sample under measurement constraints
- Subsampling approach for least squares fitting of semi-parametric accelerated failure time models to massive survival data
- Optimal subsampling for modal regression in massive data
- Emerging directions in Bayesian computation
- Optimal subsampling for the Cox proportional hazards model with massive survival data
- A review on design inspired subsampling for big data
- Approximating Partial Likelihood Estimators via Optimal Subsampling
- Deterministic subsampling for logistic regression with massive data
- Robust optimal subsampling based on weighted asymmetric least squares
- A novel residual subsampling method for skew-normal mode regression model with massive data
- Robust and efficient subsampling algorithms for massive data logistic regression
- Poisson subsampling-based estimation for growing-dimensional expectile regression in massive data
- Optimal sampling for positive only electronic health record data
- Distributed optimal subsampling for quantile regression with massive data
- Optimal Poisson subsampling for softmax regression
- A distance metric-based space-filling subsampling method for nonparametric models
- Distributed subsampling for multiplicative regression
- The COR criterion for optimal subset selection in distributed estimation
- On the inversion-free Newton's method and its applications
- A selective review on statistical methods for massive data computation: distributed computing, subsampling, and minibatch techniques
- Feature Screening for Massive Data Analysis by Subsampling
- Sampling-based estimation for massive survival data with additive hazards model
- Optimal subsampling for parametric accelerated failure time models with massive survival data
- Density Regression with Conditional Support Points
- A Subsampling Method for Regression Problems Based on Minimum Energy Criterion
- Efficient Model-Free Subsampling Method for Massive Data
- Core-elements for large-scale least squares estimation
- Optimal Subsampling via Predictive Inference
- Optimal subsampling for semi-parametric accelerated failure time models with massive survival data using a rank-based approach
- Optimal Poisson subsampling decorrelated score for high-dimensional generalized linear models
- Differential evolution variants for searching D- and A-optimal designs for nonlinear models in the bioscience
- Subsampling adaptive projection-test with mixed predictors
- D-optimal subsampling design for multiple linear regression on massive data
This page was built for publication: Optimal subsampling for large sample logistic regression
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4962448)