When does gradient descent with logistic loss find interpolating two-layer networks?
From MaRDI portal
Recommendations
- Gradient descent optimizes over-parameterized deep ReLU networks
- Gradient descent on infinitely wide neural networks: global convergence and generalization
- The implicit bias of gradient descent on separable data
- Gradient descent provably escapes saddle points in the training of shallow ReLU networks
- A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
Cites work
- Adaptive estimation of a quadratic functional by model selection.
- Benign overfitting in linear regression
- Convergence results for neural networks via electrodynamics
- Convex optimization: algorithms and complexity
- Gradient descent optimizes over-parameterized deep ReLU networks
- High-dimensional probability. An introduction with applications in data science
- High-dimensional statistics. A non-asymptotic viewpoint
- scientific article; zbMATH DE number 7370646 (Why is no real title available?)
- Introduction to algorithms.
- Just interpolate: kernel ``ridgeless regression can generalize
- Learning sparse polynomial functions
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- The implicit bias of gradient descent on separable data
Cited in
(4)- scientific article; zbMATH DE number 7625201 (Why is no real title available?)
- Binary classification of Gaussian mixtures: abundance of support vectors, benign overfitting, and regularization
- When does gradient descent with logistic loss find interpolating two-layer networks?
- On the robustness of the minimim _2 interpolator
This page was built for publication: When does gradient descent with logistic loss find interpolating two-layer networks?
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5159434)