Entropic gradient descent algorithms and wide flat minima*
From MaRDI portal
Publication:5020063
Recommendations
- Shaping the learning landscape in neural networks around wide flat minima
- Entropy-SGD: biasing gradient descent into wide valleys
- Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures
- The inverse variance-flatness relation in stochastic gradient descent is critical for finding flat minima
- Flat minima generalize for low-rank matrix recovery
Cites work
- Deep relaxation: partial differential equations for optimizing deep neural networks
- Entropy-SGD: biasing gradient descent into wide valleys
- Flat Minima
- scientific article; zbMATH DE number 149062 (Why is no real title available?)
- scientific article; zbMATH DE number 1273988 (Why is no real title available?)
- Information, Physics, and Computation
- Local entropy as a measure for sampling solutions in constraint satisfaction problems
- Shaping the learning landscape in neural networks around wide flat minima
Cited in
(6)- Deep networks on toroids: removing symmetries reveals the structure of flat regions in the landscape geometry*
- Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures
- Deep relaxation of controlled stochastic gradient descent via singular perturbations
- Visualizing high-dimensional loss landscapes with Hessian directions
- Flat minima generalize for low-rank matrix recovery
- Controlled Langevin dynamics for sampling of feedforward neural networks trained with minibatches
This page was built for publication: Entropic gradient descent algorithms and wide flat minima*
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5020063)