H2Opus: a distributed-memory multi-GPU software package for non-local operators
From MaRDI portal
Abstract: Hierarchical -matrices are asymptotically optimal representations for the discretizations of non-local operators such as those arising in integral equations or from kernel functions. Their complexity in both memory and operator application makes them particularly suited for large-scale problems. As a result, there is a need for software that provides support for distributed operations on these matrices to allow large-scale problems to be represented. In this paper, we present high-performance, distributed-memory GPU-accelerated algorithms and implementations for matrix-vector multiplication and matrix recompression of hierarchical matrices in the format. The algorithms are a new module of H2Opus, a performance-oriented package that supports a broad variety of -matrix operations on CPUs and GPUs. Performance in the distributed GPU setting is achieved by marshaling the tree data of the hierarchical matrix representation to allow batched kernels to be executed on the individual GPUs. MPI is used for inter-process communication. We optimize the communication data volume and hide much of the communication cost with local compute phases of the algorithms. Results show near-ideal scalability up to 1024 NVIDIA V100 GPUs on Summit, with performance exceeding 2.3 Tflop/s/GPU for the matrix-vector multiplication, and 670 Gflops/s/GPU for matrix compression, which involves batched QR and SVD operations. We illustrate the flexibility and efficiency of the library by solving a 2D variable diffusivity integral fractional diffusion problem with an algebraic multigrid-preconditioned Krylov solver and demonstrate scalability up to 16M degrees of freedom problems on 64 GPUs.
Recommendations
- HONEI: A collection of libraries for numerical computations targeting multiple processor architectures
- HPMaX: heterogeneous parallel matrix multiplication using CPUs and GPUs
- Optimizing the multipole-to-local operator in the fast multipole method for graphical processing units
- A direct parallel implementation of the Hoshen-Kopelman algorithm for distributed memory architectures
- HPC\(^2\) -- a fully-portable, algebra-based framework for heterogeneous computing. Application to CFD
Cites work
- A distributed-memory package for dense hierarchically semi-separable matrix computations using randomization
- A fast algorithm for simulating multiphase flows through periodic geometries of arbitrary shape
- A simple solver for the fractional Laplacian in multiple dimensions
- A spectrally accurate direct solution technique for frequency-domain scattering problems with variable media
- Algorithmic patterns for \(\mathcal {H}\)-matrices on many-core processors
- An O(N) algorithm for constructing the solution operator to 2D elliptic boundary value problems in the absence of body loads
- Construction and arithmetics of \(\mathcal H\)-matrices
- Data-sparse approximation by adaptive \({\mathcal H}^2\)-matrices
- Distributed $${{\mathcal H}^2}$$ -matrices for non-local operators
- Efficient numerical methods for non-local operators. \(\mathcal H^2\)-matrix compression, algorithms and analysis.
- Fast multipole methods for the evaluation of layer potentials with locally-corrected quadratures
- Finding structure with randomness: probabilistic algorithms for constructing approximate matrix decompositions
- H2Pack
- Hierarchical matrices: algorithms and analysis
- Hierarchical Matrix Approximations of Hessians Arising in Inverse Problems Governed by PDEs
- Hierarchical matrix operations on GPUs. Matrix-vector multiplication and compression
- High-order accurate methods for Nyström discretization of integral equations on smooth curves in the plane
- Hm-toolbox: MATLAB software for HODLR and HSS matrices
- KSPHPDDM and PCHPDDM: extending PETSc with advanced Krylov methods and robust multilevel overlapping Schwarz preconditioners
- Randomized GPU Algorithms for the Construction of Hierarchical Matrices from Matrix-Vector Operations
- Semiautomatic task graph construction for \(\mathcal{H}\)-matrix arithmetic
- Solving boundary integral problems with BEM++
- Zeta correction: a new approach to constructing corrected trapezoidal quadrature rules for singular integral operators
Cited in
(4)
Describes a project that uses
Uses Software
This page was built for publication: H2Opus: a distributed-memory multi-GPU software package for non-local operators
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2673500)