Pretest estimation in combining probability and non-probability samples
From MaRDI portal
(Redirected from Publication:6170605)
Abstract: Multiple heterogeneous data sources are becoming increasingly available for statistical analyses in the era of big data. As an important example in finite-population inference, we develop a unified framework of the test-and-pool approach to general parameter estimation by combining gold-standard probability and non-probability samples. We focus on the case when the study variable is observed in both datasets for estimating the target parameters, and each contains other auxiliary variables. Utilizing the probability design, we conduct a pretest procedure to determine the comparability of the non-probability data with the probability data and decide whether or not to leverage the non-probability data in a pooled analysis. When the probability and non-probability data are comparable, our approach combines both data for efficient estimation. Otherwise, we retain only the probability data for estimation. We also characterize the asymptotic distribution of the proposed test-and-pool estimator under a local alternative and provide a data-adaptive procedure to select the critical tuning parameters that target the smallest mean square error of the test-and-pool estimator. Lastly, to deal with the non-regularity of the test-and-pool estimator, we construct a robust confidence interval that has a good finite-sample coverage property.
Cites work
- A Simplex Method for Function Minimization
- Adaptive confidence intervals for the test error in classification
- Adjusting for Nonignorable Drop-Out Using Semiparametric Nonresponse Models
- Asymptotic results for multiple imputation
- Asymptotic Statistics
- Bias-reduced doubly robust estimation
- Big Data, Official Statistics and Some Initiatives by the Australian Bureau of Statistics
- Bootstrap methods for imputed data from regression, ratio and hot-deck imputation
- Bootstrap Sample Size in Nonregular Cases
- Calibration Estimators in Survey Sampling
- Combining multiple observational data sources to estimate causal effects
- Developments in Survey Research over the Past 60 Years: A Personal Perspective
- Doubly Robust Inference when Combining Probability and Non-Probability Samples with High Dimensional Data
- Doubly robust inference with missing data in survey sampling
- Doubly Robust Inference With Nonprobability Survey Samples
- Dynamic treatment regimes: technical challenges and applications
- Elliptical and Radial Truncation in Normal Populations
- Essential statistical inference. Theory and methods
- Estimation of Regression Coefficients When Some Regressors Are Not Always Observed
- Fixed effects, random effects or Hausman-Taylor: a pretest estimator
- scientific article; zbMATH DE number 5018147 (Why is no real title available?)
- scientific article; zbMATH DE number 3273551 (Why is no real title available?)
- Inference for nonprobability samples
- Inference for optimal dynamic treatment regimes using an adaptive m-out-of-n bootstrap scheme
- Instrumental Variables Regression with Weak Instruments
- Model assisted survey sampling.
- Modelling Overdispersion for Complex Survey Data
- Models for Nonresponse in Sample Surveys
- On making valid inferences by integrating data from surveys and other sources
- Optimal Structural Nested Models for Optimal Sequential Decisions
- Pseudo-likelihood and quasi-likelihood estimation for complex sampling schemes
- Relations between Weak and Uniform Convergence of Measures with Applications
- Sampling Statistics
- Sampling Techniques for Big Data Analysis
- Semiparametric theory and missing data.
- Statistical data integration in survey sampling: a review
- The central role of the propensity score in observational studies for causal effects
This page was built for publication: Pretest estimation in combining probability and non-probability samples
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6170605)