Flat Minima
From MaRDI portal
Recommendations
Cites work
- A Mathematical Theory of Communication
- An Information Measure for Classification
- Generalized Cross-Validation as a Method for Choosing a Good Ridge Parameter
- Modeling by shortest data description
- Smoothing noisy data with spline functions: Estimating the correct degree of smoothing by the method of generalized cross-validation
- Statistical predictor identification
Cited in
(45)- Adaptive regularization parameter selection method for enhancing generalization capability of neural networks
- Interpretable machine learning: fundamental principles and 10 grand challenges
- A spin glass model for the loss surfaces of generative adversarial networks
- Noise-induced degeneration in online learning
- Lipschitzness is all you need to tame off-policy generative adversarial imitation learning
- Optimization for deep learning: an overview
- Global optimization issues in deep network regression: an overview
- Flatness pattern recognition based on a binary tree hierarchical BP model
- Emergence of invariance and disentanglement in deep representations
- Machine learning the kinematics of spherical particles in fluid flows
- On Different Facets of Regularization Theory
- Unique sharp local minimum in \(\ell_1\)-minimization complete dictionary learning
- Minimum description length revisited
- Structure-preserving deep learning
- Hausdorff dimension, heavy tails, and generalization in neural networks*
- Entropic gradient descent algorithms and wide flat minima*
- Deep learning in target space
- Deep networks on toroids: removing symmetries reveals the structure of flat regions in the landscape geometry*
- ‘Place-cell’ emergence and learning of invariant data with restricted Boltzmann machines: breaking and dynamical restoration of continuous symmetries in the weight space
- Archetypal landscapes for deep neural networks
- The inverse variance-flatness relation in stochastic gradient descent is critical for finding flat minima
- scientific article; zbMATH DE number 7307488 (Why is no real title available?)
- Removing potential flat spots on error surface of multilayer perceptron (MLP) neural networks
- Entropy-SGD: biasing gradient descent into wide valleys
- Shaping the learning landscape in neural networks around wide flat minima
- Universal statistics of Fisher information in deep neural networks: mean field approach*
- Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures
- Prediction errors for penalized regressions based on generalized approximate message passing
- Geometric characterization of the Eyring-Kramers formula
- Lotka-Volterra model with mutations and generative adversarial networks
- Diametrical risk minimization: theory and computations
- Flat minima generalize for low-rank matrix recovery
- Neural networks generalize on low complexity data
- Homogenization of SGD in high-dimensions: exact dynamics and generalization properties
- Characterizing dynamical stability of stochastic gradient descent in overparameterized learning
- Towards a mathematical understanding of neural network-based machine learning: what we know and what we don't
- Enhancing accuracy in deep learning using random matrix theory
- Convergence of ease-controlled random reshuffling gradient algorithms under Lipschitz smoothness
- Per-example gradient regularization improves learning signals from noisy data
- Optimal control of SPDEs driven by time-space Brownian motion
- A geometric modeling of Occam's razor in deep learning
- How to explain grokking
- Loss barcode: a topological measure of escapability in loss landscapes
- Title not available (Why is no real title available?)
- Title not available (Why is no real title available?)
This page was built for publication: Flat Minima
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3123284)