Distributed coordinate descent method for learning with big data
From MaRDI portal
Abstract: In this paper we develop and analyze Hydra: HYbriD cooRdinAte descent method for solving loss minimization problems with big data. We initially partition the coordinates (features) and assign each partition to a different node of a cluster. At every iteration, each node picks a random subset of the coordinates from those it owns, independently from the other computers, and in parallel computes and applies updates to the selected coordinates based on a simple closed-form formula. We give bounds on the number of iterations sufficient to approximately solve the problem with high probability, and show how it depends on the data and on the partitioning. We perform numerical experiments with a LASSO instance described by a 3TB matrix.
Recommendations
- Parallel coordinate descent methods for big data optimization
- Efficiency of coordinate descent methods on huge-scale optimization problems
- Coordinate descent algorithms
- Accelerated, parallel, and proximal coordinate descent
- Optimization in high dimensions via accelerated, parallel, and proximal coordinate descent
Cited in
(40)- Matrix completion under interval uncertainty
- An accelerated distributed gradient method with local memory
- Parallel random block-coordinate forward-backward algorithm: a unified convergence analysis
- Primal-dual block-proximal splitting for a class of non-convex problems
- Distributed cooperative learning over time-varying random networks using a gossip-based communication protocol
- Block-proximal methods with spatially adapted acceleration
- Parallel coordinate descent methods for big data optimization
- On the convergence analysis of asynchronous SGD for solving consistent linear systems
- Manifold optimization for hybrid beamforming in dual-function radar-communication system
- Optimization in high dimensions via accelerated, parallel, and proximal coordinate descent
- Efficiency of the accelerated coordinate descent method on structured optimization problems
- On optimal probabilities in stochastic coordinate descent methods
- A Randomized Exchange Algorithm for Computing Optimal Approximate Designs of Experiments
- Accelerated, parallel, and proximal coordinate descent
- An accelerated randomized proximal coordinate gradient method and its application to regularized empirical risk minimization
- Distributed block coordinate descent for minimizing partially separable functions
- Parallel random coordinate descent method for composite minimization: convergence analysis and error bounds
- scientific article; zbMATH DE number 6982318 (Why is no real title available?)
- scientific article; zbMATH DE number 6982986 (Why is no real title available?)
- An efficient distributed learning algorithm based on effective local functional approximations
- DSCOVR: randomized primal-dual block coordinate algorithms for asynchronous distributed optimization
- A distributed block coordinate descent method for training l₁ regularized linear classifiers
- On the complexity of parallel coordinate descent
- Distributed Learning with Sparse Communications by Identification
- A class of parallel doubly stochastic algorithms for large-scale learning
- On Adaptive Sketch-and-Project for Solving Linear Systems
- Distributed Sufficient Dimension Reduction for Heterogeneous Massive Data
- Stochastic reformulations of linear systems: algorithms and convergence theory
- On convergence of distributed approximate Newton methods: globalization, sharper bounds and beyond
- A generic coordinate descent solver for non-smooth convex optimisation
- Stochastic distributed learning with gradient quantization and double-variance reduction
- Distributed statistical optimization for non-randomly stored big data with application to penalized learning
- scientific article; zbMATH DE number 7733450 (Why is no real title available?)
- More communication-efficient distributed sparse learning
- Distributed multi-agent optimisation via coordination with second-order nearest neighbours
- Accelerated methods with compression for horizontal and vertical federated learning
- Distributed learning with compressed gradient differences
- Convergence in distribution of randomized algorithms: the case of partially separable optimization
- Additive-effect assisted learning
- Parallel block coordinate descent methods with identification strategies
This page was built for publication: Distributed coordinate descent method for learning with big data
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2810888)