Landscape analysis for shallow neural networks: complete classification of critical points for affine target functions
From MaRDI portal
(Redirected from Publication:2156337)
Abstract: In this paper, we analyze the landscape of the true loss of neural networks with one hidden layer and ReLU, leaky ReLU, or quadratic activation. In all three cases, we provide a complete classification of the critical points in the case where the target function is affine and one-dimensional. In particular, we show that there exist no local maxima and clarify the structure of saddle points. Moreover, we prove that non-global local minima can only be caused by `dead' ReLU neurons. In particular, they do not appear in the case of leaky ReLU or quadratic activation. Our approach is of a combinatorial nature and builds on a careful analysis of the different types of hidden neurons that can occur.
Recommendations
- Symmetry \& critical points for a model shallow neural network
- Optimization Landscape of Neural Networks
- The global optimization geometry of shallow linear neural networks
- Non-differentiable saddle points and sub-optimal local minima exist for deep ReLU networks
- Critical points for least-squares problems involving certain analytic functions, with applications to sigmoidal nets
Cites work
- A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
- Convergence rates for the stochastic gradient descent method for non-convex objective functions
- First-order methods almost always avoid strict saddle points
- Gradient descent only converges to minimizers: non-isolated critical points and invariant regions
- Spurious valleys in one-hidden-layer neural network optimization landscapes
- Theoretical Insights Into the Optimization Landscape of Over-Parameterized Shallow Neural Networks
- Topological properties of the set of functions generated by neural networks of fixed size
Cited in
(11)- Critical points for least-squares problems involving certain analytic functions, with applications to sigmoidal nets
- The global optimization geometry of shallow linear neural networks
- Symmetry \& critical points for a model shallow neural network
- Convergence analysis for gradient flows in the training of artificial neural networks with ReLU activation
- A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions
- Critical point-finding methods reveal gradient-flat regions of deep network losses
- Revisiting Landscape Analysis in Deep Neural Networks: Eliminating Decreasing Paths to Infinity
- Optimization Landscape of Neural Networks
- Non-differentiable saddle points and sub-optimal local minima exist for deep ReLU networks
- On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
- Gradient descent provably escapes saddle points in the training of shallow ReLU networks
This page was built for publication: Landscape analysis for shallow neural networks: complete classification of critical points for affine target functions
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2156337)