Redundancy techniques for straggler mitigation in distributed optimization and learning
From MaRDI portal
Abstract: Performance of distributed optimization and learning systems is bottlenecked by "straggler" nodes and slow communication links, which significantly delay computation. We propose a distributed optimization framework where the dataset is "encoded" to have an over-complete representation with built-in redundancy, and the straggling nodes in the system are dynamically left out of the computation at every iteration, whose loss is compensated by the embedded redundancy. We show that oblivious application of several popular optimization algorithms on encoded data, including gradient descent, L-BFGS, proximal gradient under data parallelism, and coordinate descent under model parallelism, converge to either approximate or exact solutions of the original problem when stragglers are treated as erasures. These convergence results are deterministic, i.e., they establish sample path convergence for arbitrary sequences of delay patterns or distributions on the nodes, and are independent of the tail behavior of the delay distribution. We demonstrate that equiangular tight frames have desirable properties as encoding matrices, and propose efficient mechanisms for encoding large-scale data. We implement the proposed technique on Amazon EC2 clusters, and demonstrate its performance over several learning problems, including matrix factorization, LASSO, ridge regression and logistic regression, and compare the proposed method with uncoded, asynchronous, and data replication strategies.
Recommendations
- Fundamental resource trade-offs for encoded distributed optimization
- Coded Computing: Mitigating Fundamental Bottlenecks in Large-Scale Distributed Computing and Machine Learning
- Group stochastic gradient descent: a tradeoff between straggler and staleness
- A distributed flexible delay-tolerant proximal gradient algorithm
- Distributed asynchronous deterministic and stochastic gradient optimization algorithms
Cites work
- A limit theorem for the norm of random matrices
- An Asynchronous Parallel Stochastic Coordinate Descent Algorithm
- ARock: an algorithmic framework for asynchronous parallel coordinate updates
- Augmented _1 and nuclear-norm models with a globally linearly convergent algorithm
- Coded Computation Over Heterogeneous Clusters
- Complex Hadamard matrices and equiangular tight frames
- Decoding by Linear Programming
- Faster least squares approximation
- Global convergence of online limited memory BFGS
- Graph Codes for Distributed Instant Message Collection in an Arbitrary Noisy Broadcast Network
- Lower bounds on the maximum cross correlation of signals (Corresp.)
- Multi-task learning for straggler avoiding predictive job scheduling
- Near-Optimal Signal Recovery From Random Projections: Universal Encoding Strategies?
- On Orthogonal Matrices
- Orthogonal Matrices with Zero Diagonal
- Queueing with redundant requests: exact analysis
- Randomized Algorithms for Matrices and Data
- Randomized Sketches of Convex Programs With Sharp Guarantees
- Redundancy techniques for straggler mitigation in distributed optimization and learning
- Speeding Up Distributed Machine Learning Using Codes
- Steiner equiangular tight frames
- The smallest eigenvalue of a large dimensional Wishart matrix
- “Short-Dot”: Computing Large Linear Transforms Distributedly Using Coded Short Dot Products
Cited in
(6)- Multi-task learning for straggler avoiding predictive job scheduling
- A robust multi-batch L-BFGS method for machine learning
- Fundamental resource trade-offs for encoded distributed optimization
- Redundancy techniques for straggler mitigation in distributed optimization and learning
- A divide-and-conquer algorithm for distributed optimization on networks
- A code-based distributed gradient descent scheme for decentralized convex optimization
This page was built for publication: Redundancy techniques for straggler mitigation in distributed optimization and learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5381126)