Reconciling modern machine-learning practice and the classical bias-variance trade-off
From MaRDI portal
Publication:5218544
Abstract: Breakthroughs in machine learning are rapidly changing science and society, yet our fundamental understanding of this technology has lagged far behind. Indeed, one of the central tenets of the field, the bias-variance trade-off, appears to be at odds with the observed behavior of methods used in the modern machine learning practice. The bias-variance trade-off implies that a model should balance under-fitting and over-fitting: rich enough to express underlying structure in data, simple enough to avoid fitting spurious patterns. However, in the modern practice, very rich models such as neural networks are trained to exactly fit (i.e., interpolate) the data. Classically, such models would be considered over-fit, and yet they often obtain high accuracy on test data. This apparent contradiction has raised questions about the mathematical foundations of machine learning and their relevance to practitioners. In this paper, we reconcile the classical understanding and the modern practice within a unified performance curve. This "double descent" curve subsumes the textbook U-shaped bias-variance trade-off curve by showing how increasing model capacity beyond the point of interpolation results in improved performance. We provide evidence for the existence and ubiquity of double descent for a wide spectrum of models and datasets, and we posit a mechanism for its emergence. This connection between the performance and the structure of machine learning models delineates the limits of classical analyses, and has implications for both the theory and practice of machine learning.
Recommendations
- Two models of double descent for weak features
- Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
- Deep learning: a statistical viewpoint
- Learning algorithms may perform worse with increasing training set size: algorithm-data incompatibility
- The Modern Mathematics of Deep Learning
Cited in
(only showing first 100 items - show all)- Over-parametrized deep neural networks minimizing the empirical risk do not generalize well
- A generic physics-informed neural network-based constitutive model for soft biological tissues
- A selective overview of deep learning
- Linearized two-layers neural networks in high dimension
- High-dimensional dynamics of generalization error in neural networks
- Dimension independent excess risk by stochastic gradient descent
- Precise statistical analysis of classification accuracies for adversarial training
- On the robustness of minimum norm interpolators and regularized empirical risk minimizers
- AdaBoost and robust one-bit compressed sensing
- Bayesian learning via neural Schrödinger-Föllmer flows
- Understanding neural networks with reproducing kernel Banach spaces
- The interpolation phase transition in neural networks: memorization and generalization under lazy training
- A sieve stochastic gradient descent estimator for online nonparametric regression in Sobolev ellipsoids
- Deep learning for inverse problems. Abstracts from the workshop held March 7--13, 2021 (hybrid meeting)
- Surprises in high-dimensional ridgeless least squares interpolation
- Counterfactual inference with latent variable and its application in mental health care
- Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration
- Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
- Neural network training using \(\ell_1\)-regularization and bi-fidelity data
- A precise high-dimensional asymptotic theory for boosting and minimum-\(\ell_1\)-norm interpolated classifiers
- Scientific machine learning through physics-informed neural networks: where we are and what's next
- Discussion of: ``Nonparametric regression using deep neural networks with ReLU activation function
- Optimization for deep learning: an overview
- Landscape and training regimes in deep learning
- A statistician teaches deep learning
- Free dynamics of feature learning processes
- Learning algebraic models of quantum entanglement
- On the influence of over-parameterization in manifold based surrogates and deep neural operators
- On the properties of bias-variance decomposition for kNN regression
- The Random Feature Model for Input-Output Maps between Banach Spaces
- scientific article; zbMATH DE number 1934544 (Why is no real title available?)
- scientific article; zbMATH DE number 7370646 (Why is no real title available?)
- Generalization error of minimum weighted norm and kernel interpolation
- Implicit Regularization and Momentum Algorithms in Nonlinearly Parameterized Adaptive Control and Prediction
- scientific article; zbMATH DE number 7387621 (Why is no real title available?)
- A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent*
- Generalisation error in learning with random features and the hidden manifold model*
- Two models of double descent for weak features
- scientific article; zbMATH DE number 7626719 (Why is no real title available?)
- scientific article; zbMATH DE number 7626737 (Why is no real title available?)
- scientific article; zbMATH DE number 7626772 (Why is no real title available?)
- scientific article; zbMATH DE number 7625163 (Why is no real title available?)
- Learning curves of generic features maps for realistic datasets with a teacher-student model*
- Deep networks on toroids: removing symmetries reveals the structure of flat regions in the landscape geometry*
- Dimensionality Reduction, Regularization, and Generalization in Overparameterized Regressions
- Binary classification of Gaussian mixtures: abundance of support vectors, benign overfitting, and regularization
- Prevalence of neural collapse during the terminal phase of deep learning training
- Overparameterized neural networks implement associative memory
- Benign overfitting in linear regression
- The inverse variance-flatness relation in stochastic gradient descent is critical for finding flat minima
- On Transversality of Bent Hyperplane Arrangements and the Topological Expressiveness of ReLU Neural Networks
- Overparameterization and generalization error: weighted trigonometric interpolation
- Benefit of Interpolation in Nearest Neighbor Algorithms
- On the benefit of width for neural networks: disappearance of basins
- The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
- Randomization as regularization: a degrees of freedom explanation for random forest success
- What causes the test error? Going beyond bias-variance via ANOVA
- When does gradient descent with logistic loss find interpolating two-layer networks?
- scientific article; zbMATH DE number 7415108 (Why is no real title available?)
- scientific article; zbMATH DE number 7415120 (Why is no real title available?)
- Shallow neural networks for fluid flow reconstruction with limited sensors
- scientific article; zbMATH DE number 7426461 (Why is no real title available?)
- scientific article; zbMATH DE number 6193685 (Why is no real title available?)
- Scaling description of generalization with number of parameters in deep learning
- A multi-resolution theory for approximating infinite-\(p\)-zero-\(n\): transitional inference, individualized predictions, and a world without bias-variance tradeoff
- Large scale analysis of generalization error in learning using margin based classification methods
- A Unifying Tutorial on Approximate Message Passing
- For interpolating kernel machines, minimizing the norm of the ERM solution maximizes stability
- Prediction errors for penalized regressions based on generalized approximate message passing
- Mehler’s Formula, Branching Process, and Compositional Kernels of Deep Neural Networks
- Double Double Descent: On Generalization Errors in Transfer Learning between Linear Regression Tasks
- Deep learning: a statistical viewpoint
- Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
- Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
- HARFE: hard-ridge random feature expansion
- SCORE: approximating curvature information under self-concordant regularization
- Deep empirical risk minimization in finance: Looking into the future
- High dimensional binary classification under label shift: phase transition and regularization
- Large-dimensional random matrix theory and its applications in deep learning and wireless communications
- On the Inconsistency of Kernel Ridgeless Regression in Fixed Dimensions
- A Universal Trade-off Between the Model Size, Test Loss, and Training Loss of Linear Predictors
- Reliable extrapolation of deep neural operators informed by physics or sparse observations
- Re-thinking high-dimensional mathematical statistics. Abstracts from the workshop held May 15--21, 2022
- Is deep learning a useful tool for the pure mathematician?
- Also for \(k\)-means: more data does not imply better performance
- Random neural networks in the infinite width limit as Gaussian processes
- Stability of the scattering transform for deformations with minimal regularity
- High-Dimensional Analysis of Double Descent for Linear Regression with Random Projections
- On the robustness of sparse counterfactual explanations to adverse perturbations
- An instance-dependent simulation framework for learning with label noise
- On lower bounds for the bias-variance trade-off
- Benign Overfitting and Noisy Features
- Learning ability of interpolating deep convolutional neural networks
- The mathematics of artificial intelligence
- Approximate spectral decomposition of Fisher information matrix for simple ReLU networks
- Deep learning from a statistical perspective
- Symbolic regression via neural networks
- The curse of overparametrization in adversarial training: precise analysis of robust generalization for random features regression
- Optimization of random feature method in the high-precision regime
- Benign overfitting and adaptive nonparametric regression
This page was built for publication: Reconciling modern machine-learning practice and the classical bias-variance trade-off
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5218544)