Scaling description of generalization with number of parameters in deep learning
From MaRDI portal
Abstract: Supervised deep learning involves the training of neural networks with a large number of parameters. For large enough , in the so-called over-parametrized regime, one can essentially fit the training data points. Sparsity-based arguments would suggest that the generalization error increases as grows past a certain threshold . Instead, empirical studies have shown that in the over-parametrized regime, generalization error keeps decreasing with . We resolve this paradox through a new framework. We rely on the so-called Neural Tangent Kernel, which connects large neural nets to kernel methods, to show that the initialization causes finite-size random fluctuations of the neural net output function around its expectation . These affect the generalization error for classification: under natural assumptions, it decays to a plateau value in a power-law fashion . This description breaks down at a so-called jamming transition . At this threshold, we argue that diverges. This result leads to a plausible explanation for the cusp in test error known to occur at . Our results are confirmed by extensive empirical observations on the MNIST and CIFAR image datasets. Our analysis finally suggests that, given a computational envelope, the smallest generalization error is obtained using several networks of intermediate sizes, just beyond , and averaging their outputs.
Recommendations
- A jamming transition from under- to over-parametrization affects generalization in deep learning
- High-dimensional dynamics of generalization error in neural networks
- An analysis of training and generalization errors in shallow and deep networks
- Over-parametrized deep neural networks minimizing the empirical risk do not generalize well
- Deep learning: a statistical viewpoint
Cites work
- Bayesian learning for neural networks
- Comparing dynamics: deep neural networks versus glassy systems
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- Statistical mechanics of learning
- The implicit bias of gradient descent on separable data
- The simplest model of jamming
Cited in
(26)- A jamming transition from under- to over-parametrization affects generalization in deep learning
- Over-parametrized deep neural networks minimizing the empirical risk do not generalize well
- A phase transition for finding needles in nonlinear haystacks with LASSO artificial neural networks
- The generalization error of max-margin linear classifiers: benign overfitting and high dimensional asymptotics in the overparametrized regime
- High-Dimensional Analysis of Double Descent for Linear Regression with Random Projections
- Universality laws for Gaussian mixtures in generalized linear models
- The common intuition to transfer learning can win or lose: case studies for linear regression
- Harmonic analysis of network systems via kernels and their boundary realizations
- An analytic theory of shallow networks dynamics for hinge loss classification*
- Overparameterization and generalization error: weighted trigonometric interpolation
- The inverse variance-flatness relation in stochastic gradient descent is critical for finding flat minima
- Fluctuations, bias, variance and ensemble of learners: exact asymptotics for convex losses in high-dimension
- Learning sparse features can lead to overfitting in neural networks
- Redundant representations help generalization in wide neural networks
- A Generalization Gap Estimation for Overparameterized Models via the Langevin Functional Variance
- Large-dimensional random matrix theory and its applications in deep learning and wireless communications
- Triple descent and the two kinds of overfitting: where and why do they appear?*
- Limit shapes and large deviations in classical and quantum neural networks
- Normalization effects on deep neural networks
- Double Double Descent: On Generalization Errors in Transfer Learning between Linear Regression Tasks
- Geometric compression of invariant manifolds in neural networks
- Landscape and training regimes in deep learning
- Free dynamics of feature learning processes
- Normalization effects on shallow neural networks and related asymptotic expansions
- scientific article; zbMATH DE number 7415098 (Why is no real title available?)
- Surprises in high-dimensional ridgeless least squares interpolation
This page was built for publication: Scaling description of generalization with number of parameters in deep learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5856249)