Estimation of predictive performance in high-dimensional data settings using learning curves
From MaRDI portal
Abstract: In high-dimensional prediction settings, it remains challenging to reliably estimate the test performance. To address this challenge, a novel performance estimation framework is presented. This framework, called Learn2Evaluate, is based on learning curves by fitting a smooth monotone curve depicting test performance as a function of the sample size. Learn2Evaluate has several advantages compared to commonly applied performance estimation methodologies. Firstly, a learning curve offers a graphical overview of a learner. This overview assists in assessing the potential benefit of adding training samples and it provides a more complete comparison between learners than performance estimates at a fixed subsample size. Secondly, a learning curve facilitates in estimating the performance at the total sample size rather than a subsample size. Thirdly, Learn2Evaluate allows the computation of a theoretically justified and useful lower confidence bound. Furthermore, this bound may be tightened by performing a bias correction. The benefits of Learn2Evaluate are illustrated by a simulation study and applications to omics data.
Recommendations
- An imputation method for estimating the learning curve in classification problems
- Improved small-sample estimation of nonlinear cross-validated prediction metrics
- Adapting prediction error estimates for biased complexity selection in high-dimensional bootstrap samples
- Estimating prediction error in microarray classification: modifications of the 0.632+ bootstrap when \(n<p\)
- Computationally efficient confidence intervals for cross-validated area under the ROC curve estimates
Cites work
- scientific article; zbMATH DE number 3163305 (Why is no real title available?)
- scientific article; zbMATH DE number 3483405 (Why is no real title available?)
- A Limited Memory Algorithm for Bound Constrained Optimization
- A comparative study of ordinary cross-validation, v-fold cross-validation and the repeated learning-testing methods
- A fast and efficient implementation of qualitatively constrained quantile smoothing splines
- A method for constructing a confidence bound for the actual error rate of a prediction rule in high dimensions
- Calculating confidence intervals for prediction error in microarray classification using resampling
- Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach
- Computationally efficient confidence intervals for cross-validated area under the ROC curve estimates
- Estimating classification error rate: repeated cross-validation, repeated hold-out and bootstrap
- No unbiased estimator of the variance of K-fold cross-validation
- Sparse nonnegative solution of underdetermined linear equations by linear programming
- Testing the prediction error difference between 2 predictors
- The area above the ordinal dominance graph and the area below the receiver operating characteristic graph
This page was built for publication: Estimation of predictive performance in high-dimensional data settings using learning curves
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6167040)