Universal halting times in optimization and machine learning
From MaRDI portal
Abstract: The authors present empirical distributions for the halting time (measured by the number of iterations to reach a given accuracy) of optimization algorithms applied to two random systems: spin glasses and deep learning. Given an algorithm, which we take to be both the optimization routine and the form of the random landscape, the fluctuations of the halting time follow a distribution that, after centering and scaling, remains unchanged even when the distribution on the landscape is changed. We observe two qualitative classes: A Gumbel-like distribution that appears in Google searches, human decision times, the QR eigenvalue algorithm and spin glasses, and a Gaussian-like distribution that appears in conjugate gradient method, deep network with MNIST input data and deep network with random input data. This empirical evidence suggests presence of a class of distributions for which the halting time is independent of the underlying distribution under some conditions.
Recommendations
Cites work
- A neural computation model for decision-making times
- Behavior of slightly perturbed Lanczos and conjugate-gradient recurrences
- How long does it take to compute the eigenvalues of a random symmetric matrix?
- Large-scale machine learning with stochastic gradient descent
- Level-spacing distributions and the Airy kernel
- Methods of conjugate gradients for solving linear systems
- On the condition number of the critically-scaled Laguerre unitary ensemble
- On the principal components of sample covariance matrices
- Predicting the Behavior of Finite Precision Lanczos and Conjugate Gradient Computations
- Random Fields and Geometry
- Random matrices and complexity of spin glasses
- Universality for Eigenvalue Algorithms on Sample Covariance Matrices
- Universality for the Toda algorithm to compute the largest eigenvalue of a random matrix
- Universality in numerical computations with random data
- Universality of covariance matrices
Cited in
(6)- Universal statistics of incubation periods and other detection times via diffusion models
- Smoothed analysis for the conjugate gradient algorithm
- Halting time is predictable for large models: a universality property and average-case analysis
- Universality in numerical computations with random data
- Universality for Eigenvalue Algorithms on Sample Covariance Matrices
- Universality in numerical computation with random data: case studies and analytical results
This page was built for publication: Universal halting times in optimization and machine learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4605658)