Classifier technology and the illusion of progress
From MaRDI portal
Abstract: A great many tools have been developed for supervised classification, ranging from early methods such as linear discriminant analysis through to modern developments such as neural networks and support vector machines. A large number of comparative studies have been conducted in attempts to establish the relative superiority of these methods. This paper argues that these comparisons often fail to take into account important aspects of real problems, so that the apparent superiority of more sophisticated methods may be something of an illusion. In particular, simple methods typically yield performance almost as good as more sophisticated methods, to the extent that the difference in performance may be swamped by other sources of uncertainty that generally are not considered in the classical supervised classification paradigm.
Recommendations
- Supervised classification and tunnel vision
- Mining supervised classification performance studies: a meta-analytic investigation
- Classifier performance as a function of distributional complexity
- Assessing the performance of classification methods
- Does modeling lead to more accurate classification? A study of relative efficiency in linear classification
Cites work
- A conversation with Leo Breiman.
- A Probabilistic Nearest Neighbour Method for Statistical Pattern Recognition
- Direct versus indirect credit scoring classifications
- Distributed artificial intelligence meets machine learning. Learning in multi-agent environments. ECAI '96 workshop LDAIS, Budapest, Hungary, August 13, 1996, ICMAS '96 workshop LIOME, Kyoto, Japan, December 10, 1996. Selected papers
- scientific article; zbMATH DE number 5309027 (Why is no real title available?)
- scientific article; zbMATH DE number 3942804 (Why is no real title available?)
- scientific article; zbMATH DE number 1461524 (Why is no real title available?)
- scientific article; zbMATH DE number 1488345 (Why is no real title available?)
- scientific article; zbMATH DE number 2118472 (Why is no real title available?)
- scientific article; zbMATH DE number 835699 (Why is no real title available?)
- scientific article; zbMATH DE number 5046126 (Why is no real title available?)
- Mining supervised classification performance studies: a meta-analytic investigation
- Modelling consumer credit risk
- Quantitative Methods in Credit Management: A Survey
- Robust classification for imprecise environments
- Statistical modeling: The two cultures. (With comments and a rejoinder).
- Strategy, methods and solving the right problem
- Supervised classification and tunnel vision
- Very simple classification rules perform well on most commonly used datasets
Cited in
(71)- Off-the-peg and bespoke classifiers for fraud detection
- Some challenges for statistics
- Covariate-adjusted tensor classification in high dimensions
- Network linear discriminant analysis
- Quantification-oriented learning based on reliable classifiers
- On the dimension effect of regularized linear discriminant analysis
- Approximate models and robust decisions
- Measuring classifier performance: a coherent alternative to the area under the ROC curve
- Linear embedding by joint robust discriminant analysis and inter-class sparsity
- Linear components of quadratic classifiers
- High-dimensional linear discriminant analysis using nonparametric methods
- Modified hybrid discriminant analysis methods and their applications in machine learning
- Monitoring rare categories in sentiment and opinion analysis: a Milan mega event on Twitter platform
- Sparse semiparametric discriminant analysis
- On the sampling distribution of resubstitution and leave-one-out error estimators for linear classifiers
- Incorporating conditional dependence in latent class models for probabilistic record linkage: does it matter?
- Data science vs. statistics: two cultures?
- Graph-based sparse linear discriminant analysis for high-dimensional classification
- Training and assessing classification rules with imbalanced data
- Maximin effects in inhomogeneous large-scale data
- Parsimonious classification via generalized linear mixed models
- Recent developments in consumer credit risk assessment
- Interpretation of black-box predictive models
- What subject matter questions motivate the use of machine learning approaches compared to statistical models for probability prediction?
- Supervised classification for a family of Gaussian functional models
- An improved estimation in regression parameter matrix in multivariate regression model
- Finding causative genes from high-dimensional data: an appraisal of statistical and machine learning approaches
- Assessing naïve Bayes as a method for screening credit applicants
- Adapting a classification rule to local and global shift when only unlabelled data are available
- Benchmarking state-of-the-art classification algorithms for credit scoring: an update of research
- An empirical comparison of classification algorithms for mortgage default prediction: evidence from a distressed mortgage market
- Standardized partition spaces
- A partial overview of the theory of statistics with functional data
- Random survival forests models for SME credit risk measurement
- Experiment databases
- A locally adapting technique for edge detection using image segmentation
- Why do simple heuristics perform well in choices with binary attributes?
- Manifold matching: joint optimization of fidelity and commensurability
- Temporally adaptive estimation of logistic classifiers on data streams
- Shrinkage estimation for the regression parameter matrix in multivariate regression model
- Data mining and machine learning in astronomy
- Online linear and quadratic discriminant analysis with adaptive forgetting for streaming classification
- Learning optimal distributionally robust individualized treatment rules
- Testing the difference between two Kolmogorov-Smirnov values in the context of receiver operating characteristic curves
- A classification updating procedure motivated by high-content screening data
- The Dantzig discriminant analysis with high dimensional data
- The mRMR variable selection method: a comparative study for functional data
- A simple model-based approach to variable selection in classification and clustering
- Integrative genetic risk prediction using non-parametric empirical Bayes classification
- Parameter identifiability in statistical machine learning: a review
- Simple tiered classifiers
- Liquid chromatography mass spectrometry-based proteomics: biological and technological as\-pects
- On the use of double cross-validation for the combination of proteomic mass spectral data for enhanced diagnosis and prediction
- The Tangent Classifier
- A Statistical Framework for Hypothesis Testing in Real Data Comparison Studies
- Rejoinder to: Probability estimation with machine learning methods for dichotomous and multicategory outcome
- Comments on: ``Probability enhanced effective dimension reduction for classifying sparse functional data
- \(\lambda \)-perceptron: an adaptive classifier for data streams
- Bayesian synthesis: combining subjective analyses, with an application to ozone data
- A review of discriminant analysis in high dimensions
- Estimation strategies for the regression coefficient parameter matrix in multivariate multiple regression
- A tutorial on individualized treatment effect prediction from randomized trials with a binary endpoint
- Statistical comparison of classifiers through Bayesian hierarchical modelling
- A high-dimensional classifier with variable selection using mirror statistics
- Forks over knives: predictive inconsistency in criminal justice algorithmic risk assessment tools
- Personalized dynamic super learning: an application in predicting hemodiafiltration convection volumes
- High-dimensional scale invariant discriminant analysis
- Comparison between splines and fractional polynomials for multivariable model building with continuous covariates: a simulation study with continuous response
- Learning models with uniform performance via distributionally robust optimization
- Verbal autopsy methods with multiple causes of death
- Feature selection in omics prediction problems using cat scores and false nondiscovery rate control
This page was built for publication: Classifier technology and the illusion of progress
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2381761)