Reducing the effect of global communication in GMRES (m) and CG on parallel distributed memory computers
The paper studies the reduction of the communication overhead introduced by inner products when the iterative methods like CG and GMRES are parallelized. Two ways of improvement are suggested. The first way consists in assembling the results of a number of inner products collectively after grouping some orthogonalization steps. The second way uses some rearranging of the computation steps to achieve a possibility of overlapping of communication with computation. The effect of both ways of improvement is assessed by deriving computer time estimations and by running test problems.
- Enlarged Krylov subspace conjugate gradient methods for reducing communication
- Publication:4945815
- Hiding global communication latency in the GMRES algorithm on massively parallel machines
- Parallelizable restarted iterative methods for nonsymmetric linear systems. II: parallel implementation
- Minimizing synchronizations in sparse iterative solvers for distributed supercomputers
- A class of Lanczos-like algorithms implemented on parallel computers
- A Newton basis GMRES implementation
- A performance model for Krylov subspace methods on mesh-based parallel computers
- BiCGstab(l) and other hybrid Bi-CG methods
- Generalized Schwarz Splittings
- GMRES: A Generalized Minimal Residual Algorithm for Solving Nonsymmetric Linear Systems
- scientific article; zbMATH DE number 434716 (Why is no real title available?)
- scientific article; zbMATH DE number 6118211 (Why is no real title available?)
- scientific article; zbMATH DE number 687994 (Why is no real title available?)
- scientific article; zbMATH DE number 736322 (Why is no real title available?)
- scientific article; zbMATH DE number 736349 (Why is no real title available?)
- scientific article; zbMATH DE number 833745 (Why is no real title available?)
- scientific article; zbMATH DE number 861473 (Why is no real title available?)
- Incomplete block LU preconditioners on slightly overlapping subdomains for a massively parallel computer
- Methods of conjugate gradients for solving linear systems
- Partitioning Sparse Matrices with Eigenvectors of Graphs
- s-step iterative methods for symmetric linear systems
- An improved parallel hybrid bi-conjugate gradient method suitable for distributed parallel computing
- Parallelizable approximate solvers for recursions arising in preconditioning
- Varying the \(s\) in your \(s\)-step GMRES
- Incomplete block LU preconditioners on slightly overlapping subdomains for a massively parallel computer
- A parallel version of GPBi-CG method suitable for distributed parallel computing
- An improved generalized conjugate residual squared (IGCRS2) algorithm suitable for distributed parallel computing
- A parallel version of QMRCGSTAB method for large linear systems in distributed parallel environments
- Analysis and practical use of flexible biCGStab
- An improved generalized conjugate residual squared algorithm suitable for distributed parallel computing
- GMRES algorithms over 35 years
- Resolved particle simulations using the Physalis method on many GPUs
- Superboundary exchange: A technique for reducing communication in distributed implementations of iterative computations
- Minimizing synchronizations in sparse iterative solvers for distributed supercomputers
- Analysis and parallel implementation of a forced N-body problem
- SOLVING SPARSE LEAST SQUARES PROBLEMS WITH PRECONDITIONED CGLS METHOD ON PARALLEL DISTRIBUTED MEMORY COMPUTERS
- The parallel computation of the smallest eigenpair of an acoustic problem with damping
- Hiding global communication latency in the GMRES algorithm on massively parallel machines
- On the cost of iterative computations
- Improved QMRCGSTAB method in distributed parallel environments
- Analyzing the effect of local rounding error propagation on the maximal attainable accuracy of the pipelined conjugate gradient method
- Conjugate residual squared method and its improvement for non-symmetric linear systems
- A parallel nearly implicit time-stepping scheme
- Alternating Anderson-Richardson method: an efficient alternative to preconditioned Krylov methods for large, sparse linear systems
- Iterative methods for unsymmetric linear systems
- A numerically stable communication-avoiding s-step GMRES algorithm
- On the backward stability of s-step GMRES
- Variable s-step technique for planar algorithms in solving indefinite linear systems
- A parallel generalized global conjugate gradient squared algorithm for linear systems with multiple right-hand sides
- An adaptive \(s\)-step conjugate gradient algorithm with dynamic basis updating.
- An improved bi-conjugate residual algorithm suitable for distributed parallel computing
- An improved GBPi-CG algorithm suitable for distributed parallel computing
This page was built for publication: Reducing the effect of global communication in \(\text{GMRES} (m)\) and CG on parallel distributed memory computers
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1904023)