Infinite-width limit of deep linear neural networks

From MaRDI portal
Publication:6587580





This paper studies an important topic in artificial neural networks, namely a description of the training dynamics of these networks in the infinite-width limit. In recent years, this study has helped scientists to understand various aspects of deep learning such as (1) the importance of the choice of scalings/parametrization when passing to the limit since several well-behaved but fundamentally different limits can be obtained. (2) The characterization of the long-term behavior of the dynamics such as global convergence which helps understanding the learning abilities of neural nets and (3) the existence of well-posed limits. Problems of this kind are well understood for two-layer neural nets, but their are many open problems for deeper neural nets. The paper under review studies the infinite-width limit of deep linear neural networks which are initialized with random parameters. The authors study several things. (1) First they show that when the number of parameters diverges, the training dynamics converge to the dynamics obtained from a gradient descent on an infinitely wide deterministic linear neural net. (2) If the weights remain random, they obtain a precise law along the training dynamics, and prove a quantitative convergence result of the linear predictor in terms of the number of parameters. (3) The authors study the continuous time limit obtained for infinitely wide linear neural nets and show that the linear predictors of the neural net converge at an exponential rate to the minimal \(l_2\)-norm minimizer of the risk.\N\NThe paper is well written with an excellent set of references.











This page was built for publication: Infinite-width limit of deep linear neural networks

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6587580)