Parallel and distributed asynchronous adaptive stochastic gradient methods
From MaRDI portal
Abstract: Stochastic gradient methods (SGMs) are the predominant approaches to train deep learning models. The adaptive versions (e.g., Adam and AMSGrad) have been extensively used in practice, partly because they achieve faster convergence than the non-adaptive versions while incurring little overhead. On the other hand, asynchronous (async) parallel computing has exhibited significantly higher speed-up over its synchronous (sync) counterpart. Async-parallel non-adaptive SGMs have been well studied in the literature from the perspectives of both theory and practical performance. Adaptive SGMs can also be implemented without much difficulty in an async-parallel way. However, to the best of our knowledge, no theoretical result of async-parallel adaptive SGMs has been established. The difficulty for analyzing adaptive SGMs with async updates originates from the second moment term. In this paper, we propose an async-parallel adaptive SGM based on AMSGrad. We show that the proposed method inherits the convergence guarantee of AMSGrad for both convex and non-convex problems, if the staleness (also called delay) caused by asynchrony is bounded. Our convergence rate results indicate a nearly linear parallelization speed-up if , where is the staleness and is the number of iterations. The proposed method is tested on both convex and non-convex machine learning problems, and the numerical results demonstrate its clear advantages over the sync counterpart and the async-parallel nonadaptive SGM.
Recommendations
- On the parallelization upper bound for asynchronous stochastic gradients descent in non-convex optimization
- A sharp convergence rate for a model equation of the asynchronous stochastic gradient descent
- Improved asynchronous parallel optimization analysis for stochastic incremental methods
- Distributed asynchronous deterministic and stochastic gradient optimization algorithms
- Group stochastic gradient descent: a tradeoff between straggler and staleness
Cites work
- A Stochastic Approximation Method
- Acceleration of Stochastic Approximation by Averaging
- Adaptive subgradient methods for online learning and stochastic optimization
- An Asynchronous Mini-Batch Algorithm for Regularized Stochastic Optimization
- An Asynchronous Parallel Stochastic Coordinate Descent Algorithm
- ARock: an algorithmic framework for asynchronous parallel coordinate updates
- Asynchronous Gradient Push
- Distributed learning systems with first-order methods
- Improved asynchronous parallel optimization analysis for stochastic incremental methods
- On the convergence of asynchronous parallel iteration with unbounded delays
- Perturbed iterate analysis for asynchronous stochastic optimization
- Robust Stochastic Approximation Approach to Stochastic Programming
- Some aspects of parallel and distributed iterative algorithms - a survey
- Stochastic First- and Zeroth-Order Methods for Nonconvex Stochastic Programming
Cited in
(19)- A sharp convergence rate for a model equation of the asynchronous stochastic gradient descent
- Parallel stochastic gradient algorithms for large-scale matrix completion
- On the parallelization upper bound for asynchronous stochastic gradients descent in non-convex optimization
- An Asynchronous Mini-Batch Algorithm for Regularized Stochastic Optimization
- Improved asynchronous parallel optimization analysis for stochastic incremental methods
- Multitask Diffusion Adaptation Over<?Pub _newline ?>Asynchronous Networks
- Asynchronous Distributed ADMM for Large-Scale Optimization—Part II: Linear Convergence Analysis and Numerical Performance
- Stochastic modified equations for the asynchronous stochastic gradient descent
- Group stochastic gradient descent: a tradeoff between straggler and staleness
- Asynchronous variance-reduced block schemes for composite non-convex stochastic optimization: block-specific steplengths and adapted batch-sizes
- scientific article; zbMATH DE number 7307474 (Why is no real title available?)
- The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory
- An Uncertainty-Weighted Asynchronous ADMM Method for Parallel PDE Parameter Estimation
- An Asynchronous Parallel Stochastic Coordinate Descent Algorithm
- Asynchronous Gradient Push
- Distributed stochastic inertial-accelerated methods with delayed derivatives for nonconvex problems
- Improving the Transient Times for Distributed Stochastic Gradient Methods
- Asynchronous SGD with stale gradient dynamic adjustment for deep learning training
- A unified theoretical framework for the last-iterate convergence of stochastic adaptive optimization
This page was built for publication: Parallel and distributed asynchronous adaptive stochastic gradient methods
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6095736)