Benign Overfitting and Noisy Features
From MaRDI portal
Abstract: Modern machine learning often operates in the regime where the number of parameters is much higher than the number of data points, with zero training loss and yet good generalization, thereby contradicting the classical bias-variance trade-off. This extit{benign overfitting} phenomenon has recently been characterized using so called extit{double descent} curves where the risk undergoes another descent (in addition to the classical U-shaped learning curve when the number of parameters is small) as we increase the number of parameters beyond a certain threshold. In this paper, we examine the conditions under which extit{Benign Overfitting} occurs in the random feature (RF) models, i.e. in a two-layer neural network with fixed first layer weights. We adopt a new view of random feature and show that extit{benign overfitting} arises due to the noise which resides in such features (the noise may already be present in the data and propagate to the features or it may be added by the user to the features directly) and plays an important implicit regularization role in the phenomenon.
Recommendations
- Benign overfitting in linear regression
- Harmless overfitting: using denoising autoencoders in estimation of distribution algorithms
- On method overfitting
- Overfitting in linear feature extraction for classification of high-dimensional image data
- Generalization in Overparameterized Models
- Binary classification of Gaussian mixtures: abundance of support vectors, benign overfitting, and regularization
- Coarse decision making and overfitting
- Naive Learning Through Probability Overmatching
Cites work
- A mean field view of the landscape of two-layer neural networks
- A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent*
- A universal sampling method for reconstructing signals with simple Fourier transforms
- Benign overfitting in linear regression
- Convergence types and rates in generic Karhunen-Loève expansions with applications to sample path properties
- scientific article; zbMATH DE number 45848 (Why is no real title available?)
- scientific article; zbMATH DE number 3007865 (Why is no real title available?)
- scientific article; zbMATH DE number 7625163 (Why is no real title available?)
- Just interpolate: kernel ``ridgeless regression can generalize
- Local Rademacher complexities
- Mean field analysis of neural networks: a central limit theorem
- Neural tangent kernel: convergence and generalization in neural networks (invited paper)
- On the equivalence between kernel quadrature rules and random feature expansions
- Optimal rates for the regularized least-squares algorithm
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- Support Vector Machines
- Surprises in high-dimensional ridgeless least squares interpolation
- The elements of statistical learning. Data mining, inference, and prediction
- The Generalization Error of Random Features Regression: Precise Asymptotics and the Double Descent Curve
- Towards a unified analysis of random Fourier features
- Two models of double descent for weak features
This page was built for publication: Benign Overfitting and Noisy Features
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6185582)