Just interpolate: kernel ``ridgeless regression can generalize
From MaRDI portal
Publication:2196223
Abstract: In the absence of explicit regularization, Kernel "Ridgeless" Regression with nonlinear kernels has the potential to fit the training data perfectly. It has been observed empirically, however, that such interpolated solutions can still generalize well on test data. We isolate a phenomenon of implicit regularization for minimum-norm interpolated solutions which is due to a combination of high dimensionality of the input data, curvature of the kernel function, and favorable geometric properties of the data such as an eigenvalue decay of the empirical covariance and kernel matrices. In addition to deriving a data-dependent upper bound on the out-of-sample error, we present experimental evidence suggesting that the phenomenon occurs in the MNIST dataset.
Recommendations
- Generalization error of minimum weighted norm and kernel interpolation
- Surprises in high-dimensional ridgeless least squares interpolation
- Benign overfitting in linear regression
- Benefit of Interpolation in Nearest Neighbor Algorithms
- Overparameterization and generalization error: weighted trigonometric interpolation
Cites work
- 10.1162/153244303321897690
- A distribution-free theory of nonparametric regression
- An introduction to support vector machines and other kernel-based learning methods.
- Best choices for regularization parameters in learning theory: on the bias-variance problem.
- Generalized Cross-Validation as a Method for Choosing a Good Ridge Parameter
- scientific article; zbMATH DE number 45848 (Why is no real title available?)
- scientific article; zbMATH DE number 1332320 (Why is no real title available?)
- scientific article; zbMATH DE number 1950576 (Why is no real title available?)
- Kernel ridge regression
- Kernels for vector-valued functions: a review
- Learning Theory
- Model selection for regularized least-squares algorithm in learning theory
- On early stopping in gradient descent learning
- On the limit of the largest eigenvalue of the large dimensional sample covariance matrix
- Optimal rates for the regularized least-squares algorithm
- Regularization networks and support vector machines
- Scikit-learn: machine learning in Python
- The origins of kriging
- The spectrum of kernel random matrices
Cited in
(73)- Linearized two-layers neural networks in high dimension
- On the robustness of minimum norm interpolators and regularized empirical risk minimizers
- The interpolation phase transition in neural networks: memorization and generalization under lazy training
- A sieve stochastic gradient descent estimator for online nonparametric regression in Sobolev ellipsoids
- Canonical thresholding for nonsparse high-dimensional linear regression
- Surprises in high-dimensional ridgeless least squares interpolation
- Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration
- Learning from non-random data in Hilbert spaces: an optimal recovery perspective
- A precise high-dimensional asymptotic theory for boosting and minimum-\(\ell_1\)-norm interpolated classifiers
- Learning the mapping \(\mathbf{x}\mapsto \sum\limits_{i=1}^d x_i^2\): the cost of finding the needle in a haystack
- Improved complexities for stochastic conditional gradient methods under interpolation-like conditions
- Kernel approximation: from regression to interpolation
- scientific article; zbMATH DE number 7370646 (Why is no real title available?)
- Diversity sampling is an implicit regularization for kernel methods
- Generalization error of minimum weighted norm and kernel interpolation
- A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent*
- Deep neural networks, generic universal interpolation, and controlled ODEs
- scientific article; zbMATH DE number 7626719 (Why is no real title available?)
- scientific article; zbMATH DE number 7625163 (Why is no real title available?)
- Generalization error rates in kernel regression: the crossover from the noiseless to noisy regime*
- On the proliferation of support vectors in high dimensions*
- Locality defeats the curse of dimensionality in convolutional teacher–student scenarios*
- Binary classification of Gaussian mixtures: abundance of support vectors, benign overfitting, and regularization
- Theoretical issues in deep networks
- Benign overfitting in linear regression
- Overparameterization and generalization error: weighted trigonometric interpolation
- The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
- What causes the test error? Going beyond bias-variance via ANOVA
- When does gradient descent with logistic loss find interpolating two-layer networks?
- A multi-resolution theory for approximating infinite-\(p\)-zero-\(n\): transitional inference, individualized predictions, and a world without bias-variance tradeoff
- Multilevel Fine-Tuning: Closing Generalization Gaps in Approximation of Solution Maps under a Limited Budget for Training Data
- A Unifying Tutorial on Approximate Message Passing
- For interpolating kernel machines, minimizing the norm of the ERM solution maximizes stability
- Mehler’s Formula, Branching Process, and Compositional Kernels of Deep Neural Networks
- Deep learning: a statistical viewpoint
- Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
- Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
- HARFE: hard-ridge random feature expansion
- On the Inconsistency of Kernel Ridgeless Regression in Fixed Dimensions
- A Universal Trade-off Between the Model Size, Test Loss, and Training Loss of Linear Predictors
- SVRG meets AdaGrad: painless variance reduction
- Benign Overfitting and Noisy Features
- Learning ability of interpolating deep convolutional neural networks
- Tractability from overparametrization: the example of the negative perceptron
- Benign overfitting and adaptive nonparametric regression
- Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks
- Convergence analysis for over-parameterized deep learning
- New equivalences between interpolation and SVMs: kernels and structured features
- Deep networks for system identification: a survey
- Benign overfitting in time-series linear models with over-parameterization
- Solving partial differential equations with random feature models
- Classification in the high dimensional anisotropic mixture framework: a new take on robust interpolation
- DRM revisited: a complete error analysis
- The phase diagram of kernel interpolation in large dimensions
- The high-dimensional asymptotics of principal component regression
- Enhanced Response Envelope via Envelope Regularization
- The generalization error of max-margin linear classifiers: benign overfitting and high dimensional asymptotics in the overparametrized regime
- Convergence analysis of PINNs with over-parameterization
- On the robustness of the minimim _2 interpolator
- Knoop: practical enhancement of knockoff with over-parameterization for variable selection
- A geometrical viewpoint on the benign overfitting property of the minimum _2-norm interpolant estimator and its universality
- Ensemble linear interpolators: the role of ensembling
- Polynomial-based kernel reproduced gradient descent for stochastic optimization
- Learning curves for deep structured Gaussian feature models
- Six lectures on linearized neural networks
- Effectively leveraging momentum terms in stochastic line search frameworks for fast optimization of finite-sum problems
- Deep learning generalization and the convex hull of training sets
- Optimal rates of kernel ridge regression under source condition in large dimensions
- Universality of kernel random matrices and kernel regression in the quadratic regime
- Overparametrized linear regression with noisy and missing data
- Convergence conditions for stochastic line search based optimization of over-parametrized models
- Sobolev norm inconsistency of kernel interpolation
- Communication-efficient distributed estimator for generalized linear models with a diverging number of covariates
This page was built for publication: Just interpolate: kernel ``ridgeless regression can generalize
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2196223)