Infinite-width limit of deep linear neural networks
This paper studies an important topic in artificial neural networks, namely a description of the training dynamics of these networks in the infinite-width limit. In recent years, this study has helped scientists to understand various aspects of deep learning such as (1) the importance of the choice of scalings/parametrization when passing to the limit since several well-behaved but fundamentally different limits can be obtained. (2) The characterization of the long-term behavior of the dynamics such as global convergence which helps understanding the learning abilities of neural nets and (3) the existence of well-posed limits. Problems of this kind are well understood for two-layer neural nets, but their are many open problems for deeper neural nets. The paper under review studies the infinite-width limit of deep linear neural networks which are initialized with random parameters. The authors study several things. (1) First they show that when the number of parameters diverges, the training dynamics converge to the dynamics obtained from a gradient descent on an infinitely wide deterministic linear neural net. (2) If the weights remain random, they obtain a precise law along the training dynamics, and prove a quantitative convergence result of the linear predictor in terms of the number of parameters. (3) The authors study the continuous time limit obtained for infinitely wide linear neural nets and show that the linear predictors of the neural net converge at an exponential rate to the minimal \(l_2\)-norm minimizer of the risk.\N\NThe paper is well written with an excellent set of references.
- Wide neural networks of any depth evolve as linear models under gradient descent *
- Random neural networks in the infinite width limit as Gaussian processes
- Gradient descent on infinitely wide neural networks: global convergence and generalization
- Mean Field Analysis of Deep Neural Networks
- Mean-field limits of trained weights in deep learning: a dynamical systems perspective
- A mathematical theory of semantic development in deep neural networks
- A mean field view of the landscape of two-layer neural networks
- An iterative construction of solutions of the TAP equations for the Sherrington-Kirkpatrick model
- Approximation and estimation bounds for artificial neural networks
- Bayesian learning for neural networks
- Deep Linear Networks for Matrix Completion—an Infinite Depth Limit
- Disentangling feature and lazy training in deep neural networks
- Gradient descent on infinitely wide neural networks: global convergence and generalization
- Gradient flows on graphons: existence, convergence, continuity equations
- High-dimensional probability. An introduction with applications in data science
- Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
- Mean field analysis of neural networks: a law of large numbers
- Products of many large random matrices and gradients in deep neural networks
- Representations for partially exchangeable arrays of random variables
- The Continuous Formulation of Shallow Neural Networks as Wasserstein-Type Gradient Flows
- The Dynamics of Message Passing on Dense Graphs, with Applications to Compressed Sensing
- Phase diagram for two-layer ReLU neural networks at infinite-width limit
- Asymptotics of representation learning in finite Bayesian neural networks*
- Wide neural networks with bottlenecks are deep Gaussian processes
- Doubly infinite residual neural networks: a diffusion process approach
- Deep Linear Networks for Matrix Completion—an Infinite Depth Limit
- Deep stable neural networks: large-width asymptotics and convergence rates
- Random neural networks in the infinite width limit as Gaussian processes
- -Stable convergence of heavy-/light-tailed infinitely wide neural networks
- Gradient descent on infinitely wide neural networks: global convergence and generalization
- Mean-field limits of trained weights in deep learning: a dynamical systems perspective
- Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks
- Exact learning dynamics of deep linear networks with prior knowledge
- Self-consistent dynamical field theory of kernel evolution in wide neural networks
- Homotopy relaxation training algorithms for infinite-width two-layer ReLU neural networks
- A geometrical analysis of kernel ridge regression and its applications
- Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers
- First-order conditions for optimization in the Wasserstein space
- How feature learning can improve neural scaling laws
- Gradient flow equations for deep linear neural networks: a survey from a network perspective
This page was built for publication: Infinite-width limit of deep linear neural networks
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6587580)