AI without networks

From MaRDI portal





Abstract: Previously, statistical textbook wisdom has held that interpolating noisy data will generalize poorly, but recent work has shown that data interpolation schemes can generalize well. This could explain why overparameterized deep nets do not necessarily overfit. Optimal data interpolation schemes have been exhibited that achieve theoretical lower bounds for excess risk in any dimension for large data (Statistically Consistent Interpolation). These are non-parametric Nadaraya-Watson estimators with singular kernels. The recently proposed weighted interpolating nearest neighbors method (wiNN) is in this class, as is the previously studied Hilbert kernel interpolation scheme, in which the estimator has the form hatf(x)=sumiyiwi(x), where wi(x)=|xxi|d/sumj|xxj|d. This estimator is unique in being completely parameter-free. While statistical consistency was previously proven, convergence rates were not established. Here, we comprehensively study the finite sample properties of Hilbert kernel regression. We prove that the excess risk is asymptotically equivalent pointwise to sigma2(x)/ln(n) where sigma2(x) is the noise variance. We show that the excess risk of the plugin classifier is less than 2|f(x)1/2|1alpha,(1+varepsilon)alphasigmaalpha(x)(ln(n))fracalpha2, for any 0<alpha<1, where f is the regression function xmapstomathbbE[y|x]. We derive asymptotic equivalents of the moments of the weight functions wi(x) for large n, for instance for , . We derive an asymptotic equivalent for the Lagrange function and exhibit the nontrivial extrapolation properties of this estimator. We present heuristic arguments for a universal w2 power-law behavior of the probability density of the weights in the large n limit.












This page was built for publication: AI without networks

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6369604)