Learning ability of interpolating deep convolutional neural networks
From MaRDI portal
Abstract: It is frequently observed that overparameterized neural networks generalize well. Regarding such phenomena, existing theoretical work mainly devotes to linear settings or fully connected neural networks. This paper studies learning ability of an important family of deep neural networks, deep convolutional neural networks (DCNNs), under underparameterized and overparameterized settings. We establish the best learning rates of underparameterized DCNNs without parameter restrictions presented in the literature. We also show that, by adding well defined layers to an underparameterized DCNN, we can obtain some interpolating DCNNs that maintain the good learning rates of the underparameterized DCNN. This result is achieved by a novel network deepening scheme designed for DCNNs. Our work provides theoretical verification on how overfitted DCNNs generalize well.
Recommendations
- Universality of deep convolutional neural networks
- Just interpolate: kernel ``ridgeless regression can generalize
- Over-parametrized deep neural networks minimizing the empirical risk do not generalize well
- Deep learning theory of distribution regression with CNNs
- Learning a deep convolutional neural network via tensor decomposition
Cites work
- A distribution-free theory of nonparametric regression
- Approximation by Combinations of ReLU and Squared ReLU Ridge Functions With <inline-formula> <tex-math notation="LaTeX">$\ell^1$ </tex-math> </inline-formula> and <inline-formula>
- Approximation properties of a multilayered feedforward artificial neural network
- Approximation spaces of deep neural networks
- Benign overfitting in linear regression
- Deep learning
- Entropy and the combinatorial dimension
- Equivalence of approximation by convolutional neural networks and fully-connected networks
- Error bounds for approximations with deep ReLU networks
- scientific article; zbMATH DE number 51427 (Why is no real title available?)
- scientific article; zbMATH DE number 7064043 (Why is no real title available?)
- Linear processes in function spaces. Theory and applications
- Neural network approximation and estimation of classifiers with classification boundary in a Barron class
- Neural Network Learning
- Nonparametric regression using deep neural networks with ReLU activation function
- Optimal approximation with sparsely connected deep neural networks
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- Ten Lectures on Wavelets
- Theory of deep convolutional neural networks. III: Approximating radial functions
- Universal Consistency of Deep Convolutional Neural Networks
- Universality of deep convolutional neural networks
Cited in
(5)- Solving PDEs on spheres with physics-informed convolutional neural networks
- Approximation and estimation capability of vision transformers for hierarchical compositional models
- On the rates of convergence for learning with convolutional neural networks
- Optimal rates of approximation by shallow \(\operatorname{ReLU}^k\) neural networks and applications to nonparametric regression
- Two-dimensional deep ReLU CNN approximation for Korobov functions: a constructive approach
This page was built for publication: Learning ability of interpolating deep convolutional neural networks
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6185680)