Overparameterization and generalization error: weighted trigonometric interpolation
From MaRDI portal
Abstract: Motivated by surprisingly good generalization properties of learned deep neural networks in overparameterized scenarios and by the related double descent phenomenon, this paper analyzes the relation between smoothness and low generalization error in an overparameterized linear learning problem. We study a random Fourier series model, where the task is to estimate the unknown Fourier coefficients from equidistant samples. We derive exact expressions for the generalization error of both plain and weighted least squares estimators. We show precisely how a bias towards smooth interpolants, in the form of weighted trigonometric interpolation, can lead to smaller generalization error in the overparameterized regime compared to the underparameterized regime. This provides insight into the power of overparameterization, which is common in modern machine learning.
Recommendations
- Generalization error of minimum weighted norm and kernel interpolation
- An analysis of training and generalization errors in shallow and deep networks
- Just interpolate: kernel ``ridgeless regression can generalize
- High-dimensional dynamics of generalization error in neural networks
- Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
Cites work
- A mathematical introduction to compressive sensing
- Benign overfitting in linear regression
- Block-circulant matrices with circulant blocks, Weil sums, and mutually unbiased bases. II. The prime power case
- Deep double descent: where bigger models and more data hurt*
- Just interpolate: kernel ``ridgeless regression can generalize
- Linearized two-layers neural networks in high dimension
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- Scaling description of generalization with number of parameters in deep learning
- Surprises in high-dimensional ridgeless least squares interpolation
- The Generalization Error of Random Features Regression: Precise Asymptotics and the Double Descent Curve
- Two models of double descent for weak features
Cited in
(9)- Weighted neural tangent kernel: a generalized and improved network-induced kernel
- The interpolation phase transition in neural networks: memorization and generalization under lazy training
- Optimal learning
- Binary classification of Gaussian mixtures: abundance of support vectors, benign overfitting, and regularization
- Just interpolate: kernel ``ridgeless regression can generalize
- Double Double Descent: On Generalization Errors in Transfer Learning between Linear Regression Tasks
- Robust implicit regularization via weight normalization
- Multiscale estimates for the condition number of non-harmonic Fourier matrices
- Generalization error of minimum weighted norm and kernel interpolation
This page was built for publication: Overparameterization and generalization error: weighted trigonometric interpolation
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5088865)