A jamming transition from under- to over-parametrization affects generalization in deep learning
From MaRDI portal
(Redirected from Publication:5872795)
Recommendations
- Scaling description of generalization with number of parameters in deep learning
- High-dimensional dynamics of generalization error in neural networks
- Over-parametrized deep neural networks minimizing the empirical risk do not generalize well
- An analysis of training and generalization errors in shallow and deep networks
- The Modern Mathematics of Deep Learning
Cites work
Cited in
(23)- Surprises in high-dimensional ridgeless least squares interpolation
- Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
- Landscape and training regimes in deep learning
- On the stability and generalization of neural networks with VC dimension and fuzzy feature encoders
- A statistician teaches deep learning
- Geometric compression of invariant manifolds in neural networks
- Triple descent and the two kinds of overfitting: where and why do they appear?*
- Generalisation error in learning with random features and the hidden manifold model*
- Two models of double descent for weak features
- Learning curves of generic features maps for realistic datasets with a teacher-student model*
- Gradient descent dynamics and the jamming transition in infinite dimensions
- The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
- Comparing dynamics: deep neural networks versus glassy systems
- Scaling description of generalization with number of parameters in deep learning
- Double Double Descent: On Generalization Errors in Transfer Learning between Linear Regression Tasks
- Deep learning: a statistical viewpoint
- Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
- The common intuition to transfer learning can win or lose: case studies for linear regression
- Tradeoff of generalization error in unsupervised learning
- Fluctuations, bias, variance and ensemble of learners: exact asymptotics for convex losses in high-dimension
- Redundant representations help generalization in wide neural networks
- Tuning parameters of deep neural network training algorithms pays off: a computational study
- On deep learning compression through superoscillations with application to risk theory
This page was built for publication: A jamming transition from under- to over-parametrization affects generalization in deep learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5872795)