Suboptimal Local Minima Exist for Wide Neural Networks with Smooth Activations
From MaRDI portal
Recommendations
- On the benefit of width for neural networks: disappearance of basins
- Effect of depth and width on local minima in deep learning
- Non-differentiable saddle points and sub-optimal local minima exist for deep ReLU networks
- Shaping the learning landscape in neural networks around wide flat minima
- Spurious valleys in one-hidden-layer neural network optimization landscapes
Cites work
- A mean field view of the landscape of two-layer neural networks
- Gradient descent optimizes over-parameterized deep ReLU networks
- Guaranteed Matrix Completion via Non-Convex Factorization
- Learning ReLU Networks on Linearly Separable Data: Algorithm, Optimality, and Generalization
- Mean field analysis of neural networks: a law of large numbers
- The zero set of a real analytic function
- Theoretical Insights Into the Optimization Landscape of Over-Parameterized Shallow Neural Networks
Cited in
(8)- scientific article; zbMATH DE number 4211543 (Why is no real title available?)
- On the benefit of width for neural networks: disappearance of basins
- Non-attracting regions of local minima in deep and wide neural networks
- Effect of depth and width on local minima in deep learning
- Non-differentiable saddle points and sub-optimal local minima exist for deep ReLU networks
- Certifying the Absence of Spurious Local Minima at Infinity
- Tuning parameters of deep neural network training algorithms pays off: a computational study
- Calabi-Yau metrics through Grassmannian learning and Donaldson's algorithm
This page was built for publication: Suboptimal Local Minima Exist for Wide Neural Networks with Smooth Activations
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5870356)