Strategies for the vectorized block conjugate gradients method
From MaRDI portal
Abstract: Block Krylov methods have recently gained a lot of attraction. Due to their increased arithmetic intensity they offer a promising way to improve performance on modern hardware. Recently Frommer et al. presented a block Krylov framework that combines the advantages of block Krylov methods and data parallel methods. We review this framework and apply it on the Block Conjugate Gradients method,to solve linear systems with multiple right hand sides. In this course we consider challenges that occur on modern hardware, like a limited memory bandwidth, the use of SIMD instructions and the communication overhead. We present a performance model to predict the efficiency of different Block CG variants and compare these with experimental numerical results.
Recommendations
- Hardware-oriented Krylov methods for high-performance computing
- A block conjugate gradient method applied to linear systems with multiple right-hand sides
- The block preconditioned conjugate gradient method on vector computers
- Retooling the method of block conjugate gradients
- scientific article; zbMATH DE number 66103
Cites work
- A unified sparse matrix data format for efficient general sparse matrix-vector multiplication on modern processors with wide SIMD units
- Block Krylov subspace methods for functions of matrices
- Enlarged GMRES for solving linear systems with one or multiple right-hand sides
- Enlarged Krylov subspace conjugate gradient methods for reducing communication
- Retooling the method of block conjugate gradients
- The \textsc{Dune} framework: basic concepts and recent developments
- The block conjugate gradient algorithm and related methods
Cited in
(6)- Block conjugate gradient type methods for the approximation of bilinear form \(C^HA^{-1}B\)
- scientific article; zbMATH DE number 434532 (Why is no real title available?)
- Hardware-oriented Krylov methods for high-performance computing
- Adaptively restarted block Krylov subspace methods with low-synchronization skeletons
- Cache optimization and performance modeling of batched, small, and rectangular matrix multiplication on Intel, AMD, and Fujitsu processors
- Vectorized parallel in time methods for low-order discretizations with application to porous media problems
This page was built for publication: Strategies for the vectorized block conjugate gradients method
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5152828)