Optimal rates for multi-pass stochastic gradient methods
From MaRDI portal
Abstract: We analyze the learning properties of the stochastic gradient method when multiple passes over the data and mini-batches are allowed. We study how regularization properties are controlled by the step-size, the number of passes and the mini-batch size. In particular, we consider the square loss and show that for a universal step-size choice, the number of passes acts as a regularization parameter, and optimal finite sample bounds can be achieved by early-stopping. Moreover, we show that larger step-sizes are allowed when considering mini-batches. Our analysis is based on a unifying approach, encompassing both batch and stochastic gradient methods as special cases. As a byproduct, we derive optimal convergence results for batch gradient methods (even in the non-attainable cases).
Recommendations
- Stochastic gradient descent with Polyak's learning rate
- Minimizing finite sums with the stochastic average gradient
- Large-scale machine learning with stochastic gradient descent
- On the regularizing property of stochastic gradient descent
- A Markov Chain Theory Approach to Characterizing the Minimax Optimality of Stochastic Gradient Descent (for Least Squares)
Cites work
- Cross-validation based adaptation for regularization operators in learning theory
- Cutting-set methods for robust convex optimization with pessimizing oracles
- Kernel ridge vs. principal component regression: minimax bounds and the qualification of regularization operators
- Learning Bounds for Kernel Regression Using Effective Data Dimensionality
- Learning Theory
- Nonparametric stochastic approximation with large step-sizes
- On early stopping in gradient descent learning
- On regularization algorithms in learning theory
- On the Generalization Ability of On-Line Learning Algorithms
- Online gradient descent learning algorithms
- Optimal distributed online prediction using mini-batches
- Optimal rates for multi-pass stochastic gradient methods
- Optimal rates for the regularized least-squares algorithm
Cited in
(35)- Generalization properties of doubly stochastic learning algorithms
- Optimal prediction for high-dimensional functional quantile regression in reproducing kernel Hilbert spaces
- Kernel conjugate gradient methods with random projections
- Dimension independent excess risk by stochastic gradient descent
- From inexact optimization to learning via gradient concentration
- Understanding generalization error of SGD in nonconvex optimization
- Score-matching representative approach for big data analysis with generalized linear models
- Optimal rates for spectral algorithms with least-squares regression over Hilbert spaces
- Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
- A new filter‐based stochastic gradient algorithm for dual‐rate ARX models
- Optimal rates for multi-pass stochastic gradient methods
- Harder, Better, Faster, Stronger Convergence Rates for Least-Squares Regression
- On the regularizing property of stochastic gradient descent
- Convergences of regularized algorithms and stochastic gradient methods with random projections
- Graph-dependent implicit regularisation for distributed stochastic subgradient descent
- Generalization performance of multi-pass stochastic gradient descent with convex loss functions
- An analysis of stochastic variance reduced gradient for linear inverse problems *
- Regularization: From Inverse Problems to Large-Scale Machine Learning
- Stochastic gradient descent for linear inverse problems in Hilbert spaces
- On the Convergence of Stochastic Gradient Descent for Nonlinear Ill-Posed Problems
- scientific article; zbMATH DE number 7306853 (Why is no real title available?)
- Decentralized learning over a network with Nyström approximation using SGD
- Multi-index antithetic stochastic gradient algorithm
- Online regularized learning algorithm for functional data
- Pairwise learning problems with regularization networks and Nyström subsampling approach
- Differentially private SGD with random features
- High-dimensional limit of one-pass SGD on least squares
- High probability bounds for stochastic subgradient schemes with heavy tailed noise]
- Homogenization of SGD in high-dimensions: exact dynamics and generalization properties
- High-dimensional scaling limits and fluctuations of online least-squares SGD with smooth covariance
- Revisiting general source condition in learning over a Hilbert space
- Convergence rates of regularized Huber regression under weak moment conditions
- Optimal rates for functional linear regression with general regularization
- On the convergence of a data-driven regularized stochastic gradient descent for nonlinear ill-posed problems
- Learning theory of regularized Huber regression
This page was built for publication: Optimal rates for multi-pass stochastic gradient methods
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4637012)