Local case-control sampling: efficient subsampling in imbalanced data sets
From MaRDI portal
Abstract: For classification problems with significant class imbalance, subsampling can reduce computational costs at the price of inflated variance in estimating model parameters. We propose a method for subsampling efficiently for logistic regression by adjusting the class balance locally in feature space via an accept-reject scheme. Our method generalizes standard case-control sampling, using a pilot estimate to preferentially select examples whose responses are conditionally rare given their features. The biased subsampling is corrected by a post-hoc analytic adjustment to the parameters. The method is simple and requires one parallelizable scan over the full data set. Standard case-control sampling is inconsistent under model misspecification for the population risk-minimizing coefficients . By contrast, our estimator is consistent for provided that the pilot estimate is. Moreover, under correct specification and with a consistent, independent pilot estimate, our estimator has exactly twice the asymptotic variance of the full-sample MLE - even if the selected subsample comprises a miniscule fraction of the full data set, as happens when the original data are severely imbalanced. The factor of two improves to if we multiply the baseline acceptance probabilities by (and weight points with acceptance probability greater than 1), taking roughly times as many data points into the subsample. Experiments on simulated and real data show that our method can substantially outperform standard case-control subsampling.
Recommendations
- Local uncertainty sampling for large-scale multiclass logistic regression
- Surprise sampling: improving and extending the local case-control sampling
- Optimal subsampling for large sample logistic regression
- Optimal subsampling for softmax regression
- More efficient estimation for logistic regression with optimal subsamples
Cited in
(50)- Optimal subsampling for large-scale quantile regression
- Fused variable screening for massive imbalanced data
- Surprise sampling: improving and extending the local case-control sampling
- Subdata selection algorithm for linear model discrimination
- A two-stage optimal subsampling estimation for missing data problems with large-scale data
- Surface temperature monitoring in liver procurement via functional variance change-point analysis
- Local uncertainty sampling for large-scale multiclass logistic regression
- Optimal subsampling for softmax regression
- Estimating promotion effects in email marketing using a large-scale cross-classified Bayesian joint model for nested imbalanced data
- Likelihood Inference for Large Scale Stochastic Blockmodels With Covariates Based on a Divide-and-Conquer Parallelizable Algorithm With Communication
- Optimal subsampling for large sample logistic regression
- Efficient posterior sampling for high-dimensional imbalanced logistic regression
- More efficient estimation for logistic regression with optimal subsamples
- Randomized maximum-contrast selection: subagging for large-scale regression
- Optimal subsampling for large‐sample quantile regression with massive data
- Model constraints independent optimal subsampling probabilities for softmax regression
- Rather “Good In, Good Out” Than “Garbage In, Garbage Out”: A Comparison of Various Discrete Subsampling Algorithms Using COVID-19 Data Without a Response Variable
- Semi-supervised inference for case-control binary data under possibly mis-specified logistic models
- Post-selection Inference of High-dimensional Logistic Regression Under Case–Control Design
- Conditional characteristic feature screening for massive imbalanced data
- Subsampling in longitudinal models
- Unweighted estimation based on optimal sample under measurement constraints
- A review on design inspired subsampling for big data
- Deterministic subsampling for logistic regression with massive data
- Optimal Poisson subsampling for softmax regression
- A distance metric-based space-filling subsampling method for nonparametric models
- Semi-supervised inference for nonparametric logistic regression
- A Subsampling Method for Regression Problems Based on Minimum Energy Criterion
- A semiparametric method for risk prediction using integrated electronic health record data
- A two-part measurement error model to estimate participation in undeclared work and related earnings
- Efficient subsampling for high-dimensional data
- On the prediction of rare events when sampling from large data
- A synthetic subsampling and estimation procedure for imbalanced big data
- Subsampled one-step estimation for fast statistical inference
- Fast and efficient causal inference in large-scale data via subsampling and projection calibration
- Optimal surrogate-assisted sampling for cost-efficient validation of electronic health record outcomes
- Multi-resolution subsampling for linear classification with massive data
- Predicting rare events using training data from stratified sampling designs, with application to human-caused wildfire prediction
- Optimal subsampling for multinomial logistic models with big data
- A Subsampling Strategy for AIC-based Model Averaging with Generalized Linear Models
- Analyzing the dissemination of news by model averaging and subsampling
- An optimal subsampling design for large-scale Cox model with censored data
- Efficient distributed estimation for expectile regression in increasing dimensions
- Estimation and testing of expectile regression with efficient subsampling for massive data
- Model-free feature screening for massive high-dimensional imbalanced classification data via a fused inverse probability weighted absolute filter
- Optimal Poisson subsampling for multiplicative regressions with massive data
- Nearly optimal two-step Poisson sampling and empirical likelihood weighting estimation for M-estimation with big data
- Employees' well-being, work-related factors, and digitalization: a decision-tree approach for imbalanced data
- Efficient modelling of presence-only species data via local background sampling
- Matrix sketching for supervised classification with imbalanced classes
This page was built for publication: Local case-control sampling: efficient subsampling in imbalanced data sets
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q480957)