Regression with missing data, a comparison study of techniques based on random forests
From MaRDI portal
Publication:6050720
Abstract: In this paper we present the practical benefits of a new random forest algorithm to deal withmissing values in the sample. The purpose of this work is to compare the different solutionsto deal with missing values with random forests and describe our new algorithm performanceas well as its algorithmic complexity. A variety of missing value mechanisms (such as MCAR,MAR, MNAR) are considered and simulated. We study the quadratic errors and the bias ofour algorithm and compare it to the most popular missing values random forests algorithms inthe literature. In particular, we compare those techniques for both a regression and predictionpurpose. This work follows a first paper Gomez-Mendez and Joly (2020) on the consistency ofthis new algorithm.
Cites work
- A random forest guided tour
- Bagging predictors
- Fully conditional specification in multivariate imputation
- scientific article; zbMATH DE number 3860199 (Why is no real title available?)
- scientific article; zbMATH DE number 3567782 (Why is no real title available?)
- scientific article; zbMATH DE number 1294360 (Why is no real title available?)
- Impact of imputation of missing values on classification error for discrete data
- Inference and missing data
- Local Linear Forests
- Multivariate adaptive regression splines
- Random forests
- Recursive partitioning on incomplete data using surrogate decisions and multiple imputation
Cited in
(2)
This page was built for publication: Regression with missing data, a comparison study of techniques based on random forests
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6050720)