Data enriched linear regression
From MaRDI portal
Abstract: We present a linear regression method for predictions on a small data set making use of a second possibly biased data set that may be much larger. Our method fits linear regressions to the two data sets while penalizing the difference between predictions made by those two models. The resulting algorithm is a shrinkage method similar to those used in small area estimation. We find a Stein-type finding for Gaussian responses: when the model has 5 or more coefficients and 10 or more error degrees of freedom, it becomes inadmissible to use only the small data set, no matter how large the bias is. We also present both plug-in and AICc-based methods to tune our penalty parameter. Most of our results use an penalty, but we obtain formulas for penalized estimates when the model is specialized to the location setting. Ordinary Stein shrinkage provides an inadmissibility result for only 3 or more coefficients, but we find that our shrinkage method typically produces much lower squared errors in as few as 5 or 10 dimensions when the bias is small and essentially equivalent squared errors when the bias is large.
Recommendations
- Applied linear regression
- Linear regression
- Enhanced ridge regressions
- Improvement and extension of available linearized regression
- scientific article; zbMATH DE number 431728
- Linear regression with nested errors using probability-linked data
- Advanced linear modeling. Statistical learning and dependent data
Cites work
- scientific article; zbMATH DE number 3122730 (Why is no real title available?)
- scientific article; zbMATH DE number 3156765 (Why is no real title available?)
- scientific article; zbMATH DE number 3844851 (Why is no real title available?)
- scientific article; zbMATH DE number 4088699 (Why is no real title available?)
- scientific article; zbMATH DE number 3441460 (Why is no real title available?)
- scientific article; zbMATH DE number 859030 (Why is no real title available?)
- scientific article; zbMATH DE number 3388498 (Why is no real title available?)
- A Unified Approach to Regression Analysis Under Double-Sampling Designs
- Bayesian shrinkage methods for partially observed data with many predictors
- Classical F-Tests and Confidence Regions for Ridge Regression
- Combining Minimax Shrinkage Estimators
- Domain adaptation in regression
- Estimation of the mean of a multivariate normal distribution
- Exploiting Gene‐Environment Independence for Analysis of Case–Control Studies: An Empirical Bayes‐Type Shrinkage Estimator to Trade‐Off between Bias and Efficiency
- Introduction to Meta‐Analysis
- On Measuring and Correcting the Effects of Data Mining and Model Selection
- On the distribution of the largest eigenvalue in principal components analysis
- Regression and time series model selection in small samples
- Shrinkage estimators for robust and efficient inference in haplotype-based case-control studies
- Small area estimation: an appraisal. With comments and a rejoinder by the authors
- Some new developments in small area estimation
- Statistical Matching
- Stein's Estimation Rule and Its Competitors--An Empirical Bayes Approach
- The Estimation of Prediction Error
Cited in
(13)- Prediction interval transfer learning for linear regression using an empirical Bayes approach
- A comparison of some existing and novel methods for integrating historical models to improve estimation of coefficients in logistic regression
- Combining observational and experimental datasets using shrinkage estimators
- Turning the information-sharing dial: efficient inference from different data sources
- Likelihood‐Based Inference for the Finite Population Mean with Post‐Stratification Information Under Non‐Ignorable Non‐Response
- Improving estimation and prediction in linear regression incorporating external information from an established reduced model
- Prediction-based regularization using data augmented regression
- Adaptive and robust multi-task learning
- Improved linear regression prediction by transfer learning
- Transfer Learning under High-dimensional Generalized Linear Models
- Transfer learning for error-contaminated Poisson regression models
- Methods for combining observational and experimental causal estimates: a review
- Robust transfer learning with unreliable source data
This page was built for publication: Data enriched linear regression
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2346524)