Normalization effects on deep neural networks
From MaRDI portal
Abstract: We study the effect of normalization on the layers of deep neural networks of feed-forward type. A given layer with hidden units is allowed to be normalized by with and we study the effect of the choice of the on the statistical behavior of the neural network's output (such as variance) as well as on the test accuracy on the MNIST data set. We find that in terms of variance of the neural network's output and test accuracy the best choice is to choose the 's to be equal to one, which is the mean-field scaling. We also find that this is particularly true for the outer layer, in that the neural network's behavior is more sensitive in the scaling of the outer layer as opposed to the scaling of the inner layers. The mechanism for the mathematical analysis is an asymptotic expansion for the neural network's output. An important practical consequence of the analysis is that it provides a systematic and mathematically informed way to choose the learning rate hyperparameters. Such a choice guarantees that the neural network behaves in a statistically robust way as the grow to infinity.
Recommendations
- Normalization effects on shallow neural networks and related asymptotic expansions
- On the effect of the activation function on the distribution of hidden nodes in a deep network
- Mean Field Analysis of Deep Neural Networks
- Theoretical issues in deep networks
- A correspondence between normalization strategies in artificial and biological neural networks
Cites work
- A mean field view of the landscape of two-layer neural networks
- Approximation and estimation bounds for artificial neural networks
- Asymptotics of Reinforcement Learning with Neural Networks
- Deep learning
- DGM: a deep learning algorithm for solving partial differential equations
- Gradient descent optimizes over-parameterized deep ReLU networks
- scientific article; zbMATH DE number 3951715 (Why is no real title available?)
- scientific article; zbMATH DE number 1972910 (Why is no real title available?)
- Large deviations and mean-field theory for asymmetric random recurrent neural networks
- Machine learning strategies for systems with invariance properties
- Mean Field Analysis of Deep Neural Networks
- Mean field analysis of neural networks: a central limit theorem
- Mean field analysis of neural networks: a law of large numbers
- Multilayer feedforward networks are universal approximators
- Nonlinearity creates linear independence
- Normalization effects on shallow neural networks and related asymptotic expansions
- Reynolds averaged turbulence modelling using deep neural networks with embedded invariance
- Scaling description of generalization with number of parameters in deep learning
- Universal features of price formation in financial markets: perspectives from deep learning
Cited in
(1)
This page was built for publication: Normalization effects on deep neural networks
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6194477)