Near-Minimax Optimal Estimation With Shallow ReLU Neural Networks
From MaRDI portal
Abstract: We study the problem of estimating an unknown function from noisy data using shallow ReLU neural networks. The estimators we study minimize the sum of squared data-fitting errors plus a regularization term proportional to the squared Euclidean norm of the network weights. This minimization corresponds to the common approach of training a neural network with weight decay. We quantify the performance (mean-squared error) of these neural network estimators when the data-generating function belongs to the second-order Radon-domain bounded variation space. This space of functions was recently proposed as the natural function space associated with shallow ReLU neural networks. We derive a minimax lower bound for the estimation problem for this function space and show that the neural network estimators are minimax optimal up to logarithmic factors. This minimax rate is immune to the curse of dimensionality. We quantify an explicit gap between neural networks and linear methods (which include kernel methods) by deriving a linear minimax lower bound for the estimation problem, showing that linear methods necessarily suffer the curse of dimensionality in this function space. As a result, this paper sheds light on the phenomenon that neural networks seem to break the curse of dimensionality.
Cited in
(13)- Distributional extension and invertibility of the k-plane transform and its dual
- MARS via lasso
- Weighted variation spaces and approximation by shallow ReLU networks
- Compositional function spaces for deep learning
- Optimal rates of approximation by shallow \(\operatorname{ReLU}^k\) neural networks and applications to nonparametric regression
- Spectral Barron space for deep neural network approximation
- ReLU neural networks with linear layers are biased towards single- and multi-index models
- Integral probability metrics meet neural networks: the Radon-Kolmogorov-Smirnov test
- Random ReLU neural networks as non-Gaussian processes
- Function-space optimality of neural architectures with multivariate nonlinearities
- Dimension-independent learning rates for high-dimensional classification problems
- Stable learning using spiking neural networks equipped with affine encoders and decoders
- Sobolev norm inconsistency of kernel interpolation
This page was built for publication: Near-Minimax Optimal Estimation With Shallow ReLU Neural Networks
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6193583)