Stochastic gradient descent with Polyak's learning rate
From MaRDI portal
Publication:1983178
Abstract: Stochastic gradient descent (SGD) for strongly convex functions converges at the rate . However, achieving good results in practice requires tuning the parameters (for example the learning rate) of the algorithm. In this paper we propose a generalization of the Polyak step size, used for subgradient methods, to Stochastic gradient descent. We prove a non-asymptotic convergence at the rate with a rate constant which can be better than the corresponding rate constant for optimally scheduled SGD. We demonstrate that the method is effective in practice, and on convex optimization problems and on training deep neural networks, and compare to the theoretical rate.
Recommendations
- Lower error bounds for the stochastic gradient descent optimization algorithm: sharp convergence rates for slowly and fast decaying learning rates
- Convergence rates for the stochastic gradient descent method for non-convex objective functions
- Minimizing finite sums with the stochastic average gradient
- New Convergence Aspects of Stochastic Gradient Algorithms
- Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
Cites work
- A Stochastic Approximation Method
- Adaptive subgradient methods for online learning and stochastic optimization
- First-order methods in optimization
- Information-Based Complexity, Feedback and Dynamics in Convex Programming
- Information-Theoretic Lower Bounds on the Oracle Complexity of Stochastic Convex Optimization
- Introduction to continuous optimization
- Introductory lectures on convex optimization. A basic course.
- Laplacian smoothing stochastic gradient Markov chain Monte Carlo
- Optimization methods for large-scale machine learning
- Some methods of speeding up the convergence of iteration methods
- Variable target value subgradient method
Cited in
(36)- Convergence of stochastic gradient descent in deep neural network
- Stopping criteria for, and strong convergence of, stochastic gradient descent on Bottou-Curtis-Nocedal functions
- An adaptive Polyak heavy-ball method
- Laplacian smoothing gradient descent
- Bridging the gap between constant step size stochastic gradient descent and Markov chains
- Why random reshuffling beats stochastic gradient descent
- Accelerating deep neural network training with inconsistent stochastic gradient descent
- Lower error bounds for the stochastic gradient descent optimization algorithm: sharp convergence rates for slowly and fast decaying learning rates
- On the linear convergence of the stochastic gradient method with constant step-size
- Learning rate adaptation in stochastic gradient descent.
- Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
- Optimal rates for multi-pass stochastic gradient methods
- Convergence rates for the stochastic gradient descent method for non-convex objective functions
- Making the last iterate of SGD information theoretically optimal
- Generalization performance of multi-pass stochastic gradient descent with convex loss functions
- scientific article; zbMATH DE number 7625168 (Why is no real title available?)
- Scheduled restart momentum for accelerated stochastic gradient descent
- The stochastic delta rule: faster and more accurate deep learning through adaptive weight noise
- Stability and optimization error of stochastic gradient descent for pairwise learning
- The error-feedback framework: SGD with delayed gradients
- An adaptive gradient method with energy and momentum
- Adaptive moment estimation for universal portfolio selection strategy
- Stochastic gradient descent: where optimization meets machine learning
- Stability analysis of stochastic gradient descent for homogeneous neural networks and linear classifiers
- A novel stepsize for gradient descent method
- Greedy randomized sampling nonlinear Kaczmarz methods
- A fast non-monotone line search for stochastic gradient descent
- The accelerated tensor Kaczmarz algorithm with adaptive parameters for solving tensor systems
- A stochastic gradient method with variance control and variable learning rate for deep learning
- Non-ergodic linear convergence property of the delayed gradient descent under the strongly convexity and the Polyak-Łojasiewicz condition
- Stochastic algorithms with geometric step decay converge linearly on sharp functions
- Bolstering stochastic gradient descent with model building
- Optimized convergence of stochastic gradient descent by weighted averaging
- Polyak minorant method for convex optimization
- RTC-fPINNs: respecting temporal causality fPINNs method for time fractional fourth order partial differential equations
- The equivalence of the randomized extended Gauss-Seidel and randomized extended Kaczmarz methods
Describes a project that uses
Uses Software
This page was built for publication: Stochastic gradient descent with Polyak's learning rate
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1983178)