Test set sizing via random matrix theory
From MaRDI portal
Abstract: This paper uses techniques from Random Matrix Theory to find the ideal training-testing data split for a simple linear regression with m data points, each an independent n-dimensional multivariate Gaussian. It defines "ideal" as satisfying the integrity metric, i.e. the empirical model error is the actual measurement noise, and thus fairly reflects the value or lack of same of the model. This paper is the first to solve for the training and test size for any model in a way that is truly optimal. The number of data points in the training set is the root of a quartic polynomial Theorem 1 derives which depends only on m and n; the covariance matrix of the multivariate Gaussian, the true model parameters, and the true measurement noise drop out of the calculations. The critical mathematical difficulties were realizing that the problems herein were discussed in the context of the Jacobi Ensemble, a probability distribution describing the eigenvalues of a known random matrix model, and evaluating a new integral in the style of Selberg and Aomoto. Mathematical results are supported with thorough computational evidence. This paper is a step towards automatic choices of training/test set sizes in machine learning.
Cites work
- A matrix model for the β-Jacobi ensemble
- Asymptotics of Selberg-like integrals by lattice path counting
- Characteristic vectors of bordered matrices with infinite dimensions
- scientific article; zbMATH DE number 3886886 (Why is no real title available?)
- scientific article; zbMATH DE number 1231230 (Why is no real title available?)
- scientific article; zbMATH DE number 2177280 (Why is no real title available?)
- scientific article; zbMATH DE number 3244317 (Why is no real title available?)
- scientific article; zbMATH DE number 3107108 (Why is no real title available?)
- Joint moments of a characteristic polynomial and its derivative for the circular \(\beta \)-ensemble
- Matrix models for beta ensembles
- Moments of the eigenvalue densities and of the secular coefficients of \(\beta\)-ensembles
- Moments of the position of the maximum for GUE characteristic polynomials and for log-correlated Gaussian processes
- ON THE COMPLEX SELBERG INTEGRAL
- Optimality of training/test size and resampling effectiveness in cross-validation
- The strong limits of random matrix spectra for sample matrices of independent elements
This page was built for publication: Test set sizing via random matrix theory
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6130677)