Communication-efficient distributed optimization of self-concordant empirical loss
From MaRDI portal
Abstract: We consider distributed convex optimization problems originated from sample average approximation of stochastic optimization, or empirical risk minimization in machine learning. We assume that each machine in the distributed computing system has access to a local empirical loss function, constructed with i.i.d. data sampled from a common distribution. We propose a communication-efficient distributed algorithm to minimize the overall empirical loss, which is the average of the local empirical losses. The algorithm is based on an inexact damped Newton method, where the inexact Newton steps are computed by a distributed preconditioned conjugate gradient method. We analyze its iteration complexity and communication efficiency for minimizing self-concordant empirical loss functions, and discuss the results for distributed ridge regression, logistic regression and binary classification with a smoothed hinge loss. In a standard setting for supervised learning, the required number of communication rounds of the algorithm does not increase with the sample size, and only grows slowly with the number of machines.
Recommendations
- Communication-efficient algorithms for statistical optimization
- An efficient distributed learning algorithm based on effective local functional approximations
- A general distributed dual coordinate optimization framework for regularized loss minimization
- Efficient protocols for distributed classification and optimization
- Distributed block-diagonal approximation methods for regularized empirical risk minimization
Cited in
(26)- Distributed optimization and statistical learning for large-scale penalized expectile regression
- Distributed estimation in heterogeneous reduced rank regression: with application to order determination in sufficient dimension reduction
- Hybrid MPI/OpenMP parallel asynchronous distributed alternating direction method of multipliers
- Communication-efficient algorithms for statistical optimization
- Efficient protocols for distributed classification and optimization
- An efficient distributed learning algorithm based on effective local functional approximations
- Distributed stochastic variance reduced gradient methods by sampling extra data with replacement
- GADMM: fast and communication efficient framework for distributed machine learning
- Distributed Optimization Based on Gradient Tracking Revisited: Enhancing Convergence Rate via Surrogation
- Distributed learning systems with first-order methods
- Randomized block proximal damped Newton method for composite self-concordant minimization
- An accelerated communication-efficient primal-dual optimization framework for structured machine learning
- Distributed Sparse Composite Quantile Regression in Ultrahigh Dimensions
- Hyperfast second-order local solvers for efficient statistically preconditioned distributed optimization
- Preconditioning meets biased compression for efficient distributed optimization
- SHED: a Newton-type algorithm for federated learning based on incremental Hessian eigenvector sharing
- scientific article; zbMATH DE number 7800967 (Why is no real title available?)
- Communication-efficient distributed cubic Newton with compressed lazy Hessian
- Distributed accelerated gradient methods with restart under quadratic growth condition
- Adaptive consensus: a network pruning approach for decentralized optimization
- Balancing communication and computation in gradient tracking algorithms for decentralized optimization
- Distributed robust estimation and inference with contaminated data
- FLECS: a federated learning second-order framework via compression and sketching
- Efficient distributed optimization for large-scale high-dimensional sparse penalized Huber regression
- Improved global performance guarantees of second-order methods in convex minimization
- Distributed block-diagonal approximation methods for regularized empirical risk minimization
This page was built for publication: Communication-efficient distributed optimization of self-concordant empirical loss
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2415209)