Orthogonal subsampling for big data linear regression
From MaRDI portal
Publication:2247473
Abstract: The dramatic growth of big datasets presents a new challenge to data storage and analysis. Data reduction, or subsampling, that extracts useful information from datasets is a crucial step in big data analysis. We propose an orthogonal subsampling (OSS) approach for big data with a focus on linear regression models. The approach is inspired by the fact that an orthogonal array of two levels provides the best experimental design for linear regression models in the sense that it minimizes the average variance of the estimated parameters and provides the best predictions. The merits of OSS are three-fold: (i) it is easy to implement and fast; (ii) it is suitable for distributed parallel computing and ensures the subsamples selected in different batches have no common data points; and (iii) it outperforms existing methods in minimizing the mean squared errors of the estimated parameters and maximizing the efficiencies of the selected subsamples. Theoretical results and extensive numerical results show that the OSS approach is superior to existing subsampling approaches. It is also more robust to the presence of interactions among covariates and, when they do exist, OSS provides more precise estimates of the interaction effects than existing methods. The advantages of OSS are also illustrated through analysis of real data.
Recommendations
- Information-Based Optimal Subdata Selection for Big Data Linear Regression
- Optimal subsampling algorithms for big data regressions
- Distributed subdata selection for big data via sampling-based approach
- Optimal subsampling for large-scale quantile regression
- Optimal subsampling for linear quantile regression models
Cites work
- A review of some exchange algorithms for constructing discrete D-optimal designs
- Columnwise-Pairwise Algorithms with Applications to the Construction of Supersaturated Designs
- Construction of supersaturated designs through partially aliased interactions
- Divide-and-conquer information-based optimal subdata selection algorithm
- Exchange algorithms for constructing large spatial designs
- Fast approximation of matrix coherence and statistical leverage
- scientific article; zbMATH DE number 3176492 (Why is no real title available?)
- scientific article; zbMATH DE number 1313654 (Why is no real title available?)
- scientific article; zbMATH DE number 1089164 (Why is no real title available?)
- scientific article; zbMATH DE number 2015211 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- scientific article; zbMATH DE number 2231192 (Why is no real title available?)
- scientific article; zbMATH DE number 3067118 (Why is no real title available?)
- Information-Based Optimal Subdata Selection for Big Data Linear Regression
- On computationally tractable selection of experiments in measurement-constrained regression models
- Optimal Bayesian design applied to logistic regression experiments
- Orthogonal arrays with variable numbers of symbols
- Orthogonal arrays. Theory and applications
- Recent developments in nonregular fractional factorial designs
- Regularization and Variable Selection Via the Elastic Net
- The Dantzig selector: statistical estimation when \(p\) is much larger than \(n\). (With discussions and rejoinder).
Cited in
(40)- Optimal subsampling for least absolute relative error estimators with massive data
- Model-free global likelihood subsampling for massive data
- Linear operator‐based statistical analysis: A useful paradigm for big data
- Information-Based Optimal Subdata Selection for Big Data Linear Regression
- Information-based optimal subdata selection for non-linear models
- Optimal subsampling design for polynomial regression in one covariate
- Predictive Subdata Selection for Computer Models
- Subdata selection based on orthogonal array for big data
- Optimal sampling designs for multidimensional streaming time series with application to power grid sensor data
- A review on design inspired subsampling for big data
- Deterministic subsampling for logistic regression with massive data
- Robust optimal subsampling based on weighted asymmetric least squares
- Poisson subsampling-based estimation for growing-dimensional expectile regression in massive data
- A distance metric-based space-filling subsampling method for nonparametric models
- A Subsampling Method for Regression Problems Based on Minimum Energy Criterion
- Core-elements for large-scale least squares estimation
- Differential evolution variants for searching D- and A-optimal designs for nonlinear models in the bioscience
- D-optimal subsampling design for multiple linear regression on massive data
- Efficient subsampling for high-dimensional data
- Big data subsampling: a review
- Optimal Subsampling for Data Streams with Measurement Constrained Categorical Responses
- Orthogonal arrays: a review
- Multi-resolution subsampling for linear classification with massive data
- Optimal distributed Poisson subsampling for modal regression with massive data
- A Review of Design of Experiments Courses Offered to Undergraduate Students at American Universities
- Subsampling for big data linear models with measurement errors
- Big Data Model Building Using Dimension Reduction and Sample Selection
- Independence-Encouraging Subsampling for Nonparametric Additive Models
- Nonparametric Additive Models for Billion Observations
- Group-Orthogonal Subsampling for Hierarchical Data Based on Linear Mixed Models
- Supervised Stratified Subsampling for Predictive Analytics
- On the selection of optimal subdata for big data regression based on leverage scores
- Optimal subsampling for generalized additive models on large-scale datasets
- Information-based optimal subdata selection for clusterwise linear regression
- Efficient subsampling for exponential family models
- Poisson subsampling for large-scale functional data analysis
- Optimal Poisson subsampling for multiplicative regressions with massive data
- Balanced subsampling for big data with categorical predictors
- Distributed information-based optimal sub-data selection algorithm for big data logistic regression
- Distributed subdata selection for big data via sampling-based approach
This page was built for publication: Orthogonal subsampling for big data linear regression
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2247473)