Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
From MaRDI portal
adaptive estimationalgorithmic approximationgradient flow dynamicsinterpolationneural networksrepresentation learningreproducing kernel Hilbert space
Hilbert spaces with reproducing kernels (= (proper) functional Hilbert spaces, including de Branges-Rovnyak and other structured spaces) (46E22) Nonparametric estimation (62G05) Learning and adaptive systems in artificial intelligence (68T05) Approximation algorithms (68W25) Applications of mathematical programming (90C90)
Abstract: Consider the problem: given the data pair drawn from a population with , specify a neural network model and run gradient flow on the weights over time until reaching any stationarity. How does , the function computed by the neural network at time , relate to , in terms of approximation and representation? What are the provable benefits of the adaptive representation by neural networks compared to the pre-specified fixed basis representation in the classical nonparametric literature? We answer the above questions via a dynamic reproducing kernel Hilbert space (RKHS) approach indexed by the training process of neural networks. Firstly, we show that when reaching any local stationarity, gradient flow learns an adaptive RKHS representation and performs the global least-squares projection onto the adaptive RKHS, simultaneously. Secondly, we prove that as the RKHS is data-adaptive and task-specific, the residual for lies in a subspace that is potentially much smaller than the orthogonal complement of the RKHS. The result formalizes the representation and approximation benefits of neural networks. Lastly, we show that the neural network function computed by gradient flow converges to the kernel ridgeless regression with an adaptive kernel, in the limit of vanishing regularization. The adaptive kernel viewpoint provides new angles of studying the approximation, representation, generalization, and optimization advantages of neural networks.
Recommendations
- Understanding neural networks with reproducing kernel Banach spaces
- A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
- Expressivity of Deep Neural Networks
- Banach space representer theorems for neural networks and ridge splines
- Training neural networks with noisy data as an ill-posed problem
Cites work
- A mean field view of the landscape of two-layer neural networks
- A simple lemma on greedy approximation in Hilbert space and convergence rates for projection pursuit regression and neural network training
- Approximation and learning by greedy algorithms
- Approximation by superpositions of a sigmoidal function
- Breaking the curse of dimensionality with convex neural networks
- scientific article; zbMATH DE number 45848 (Why is no real title available?)
- scientific article; zbMATH DE number 1332320 (Why is no real title available?)
- scientific article; zbMATH DE number 2149779 (Why is no real title available?)
- Just interpolate: kernel ``ridgeless regression can generalize
- Linearized two-layers neural networks in high dimension
- Mean field analysis of neural networks: a central limit theorem
- Multilayer feedforward networks are universal approximators
- Neural Network Learning
- Nonparametric maximum likelihood estimation by the method of sieves
- Optimal rates of convergence for nonparametric estimators
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- The Variational Formulation of the Fokker--Planck Equation
Cited in
(9)- Deep learning for the partially linear Cox model
- A precise high-dimensional asymptotic theory for boosting and minimum-\(\ell_1\)-norm interpolated classifiers
- Geometric compression of invariant manifolds in neural networks
- When do neural networks outperform kernel methods?*
- On the benefit of width for neural networks: disappearance of basins
- On Kernel Method–Based Connectionist Models and Supervised Deep Learning Without Backpropagation
- Mehler’s Formula, Branching Process, and Compositional Kernels of Deep Neural Networks
- Weighted neural tangent kernel: a generalized and improved network-induced kernel
- Quantitative CLTs in deep neural networks
This page was built for publication: Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6044638)