Integrating Multisource Block-Wise Missing Data in Model Selection
From MaRDI portal
Abstract: For multi-source data, blocks of variable information from certain sources are likely missing. Existing methods for handling missing data do not take structures of block-wise missing data into consideration. In this paper, we propose a Multiple Block-wise Imputation (MBI) approach, which incorporates imputations based on both complete and incomplete observations. Specifically, for a given missing pattern group, the imputations in MBI incorporate more samples from groups with fewer observed variables in addition to the group with complete observations. We propose to construct estimating equations based on all available information, and optimally integrate informative estimating functions to achieve efficient estimators. We show that the proposed method has estimation and model selection consistency under both fixed-dimensional and high-dimensional settings. Moreover, the proposed estimator is asymptotically more efficient than the estimator based on a single imputation from complete observations only. In addition, the proposed method is not restricted to missing completely at random. Numerical studies and ADNI data application confirm that the proposed method outperforms existing variable selection methods under various missing mechanisms.
Recommendations
- Imputed factor regression for high-dimensional block-wise missing data
- Multinomial logistic factor regression for multi-source functional block-wise missing data
- Variable selection for multiply-imputed data with penalized generalized estimating equations
- Variable selection techniques after multiple imputation in high-dimensional data
- Variable selection for regression models with missing data
Cites work
- A Generalization of Sampling Without Replacement From a Finite Universe
- A Nonlinear Conjugate Gradient Method with a Strong Global Convergence Property
- A penalized EM algorithm incorporating missing data mechanism for Gaussian parameter estimation
- Adaptive Lasso for sparse high-dimensional regression models
- Bayesian sensitivity analysis of statistical models with missing data
- Consistency and asymptotic normality of the maximum likelihood estimator in generalized linear models
- Efficient estimation for longitudinal data by combining large-dimensional moment conditions
- Endogeneity in high dimensions
- Estimating the dimension of a model
- scientific article; zbMATH DE number 5957408 (Why is no real title available?)
- scientific article; zbMATH DE number 2140075 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- Large Sample Properties of Generalized Method of Moments Estimators
- LASSO-TYPE GMM ESTIMATOR
- Missing Covariates in Generalized Linear Models When the Missing Data Mechanism is Non-ignorable
- Model Selection and Estimation in Regression with Grouped Variables
- Nonconcave Penalized Likelihood With NP-Dimensionality
- On Inverse Probability Weighting for Nonmonotone Missing at Random Data
- Optimal sparse linear prediction for block-missing multi-modality data without imputation
- Rate of Convergence of Several Conjugate Gradient Algorithms
- Sure independence screening for ultrahigh dimensional feature space. With discussion and authors' reply
- The Adaptive Lasso and Its Oracle Properties
- The sparsity and bias of the LASSO selection in high-dimensional linear regression
- Variable selection and prediction with incomplete high-dimensional data
- Variable selection models based on multiple imputation with an application for predicting median effective dose and maximum effect
- Variable Selection via Nonconcave Penalized Likelihood and its Oracle Properties
Cited in
(23)- Mallows model averaging with effective model size in fragmentary data prediction
- Imputation for multisource data with comparison and assessment techniques
- Imputed factor regression for high-dimensional block-wise missing data
- Optimal sparse linear prediction for block-missing multi-modality data without imputation
- Model averaging for generalized linear models in fragmentary data prediction
- Weighted multiple blockwise imputation method for high-dimensional regression with blockwise missing data
- BlockMissingData
- Variable selection for high‐dimensional generalized linear model with block‐missing data
- Multinomial logistic factor regression for multi-source functional block-wise missing data
- Orthogonalized Kernel Debiased Machine Learning for Multimodal Data Analysis
- Penalized estimating equations for generalized linear models with multiple imputation
- Optimal Integrating Learning for Split Questionnaire Design Type Data
- Regularized Buckley-James method for right-censored outcomes with block-missing multimodal covariates
- Prediction approaches for partly missing multi-omics covariate data: a literature review and an empirical comparison study
- Jackknife model averaging for linear regression models with missing responses
- Statistical inference for high-dimensional linear regression with blockwise missing data
- Component-based regression for hybrid data
- Time-varying mediation analysis for incomplete data with application to DNA methylation study for PTSD
- Multiply robust estimation of causal effects using linked data
- Online updating estimation for heterogenous updating generalized linear models with streaming datasets
- Additive-effect assisted learning
- A Bayesian approach for selecting relevant external data (BASE): application to a study of long-term outcomes in a hemophilia gene therapy trial (HOPE-B)
- Prescriptive modality selection for classification and regression
This page was built for publication: Integrating Multisource Block-Wise Missing Data in Model Selection
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5881971)