GPU-acceleration of the ELPA2 distributed eigensolver for dense symmetric and Hermitian eigenproblems
From MaRDI portal
(Redirected from Publication:6159210)
Abstract: The solution of eigenproblems is often a key computational bottleneck that limits the tractable system size of numerical algorithms, among them electronic structure theory in chemistry and in condensed matter physics. Large eigenproblems can easily exceed the capacity of a single compute node, thus must be solved on distributed-memory parallel computers. We here present GPU-oriented optimizations of the ELPA two-stage tridiagonalization eigensolver (ELPA2). On top of cuBLAS-based GPU offloading, we add a CUDA kernel to speed up the back-transformation of eigenvectors, which can be the computationally most expensive part of the two-stage tridiagonalization algorithm. We benchmark the performance of this GPU-accelerated eigensolver on two hybrid CPU-GPU architectures, namely a compute cluster based on Intel Xeon Gold CPUs and NVIDIA Volta GPUs, and the Summit supercomputer based on IBM POWER9 CPUs and NVIDIA Volta GPUs. Consistent with previous benchmarks on CPU-only architectures, the GPU-accelerated two-stage solver exhibits a parallel performance superior to the one-stage counterpart. Finally, we demonstrate the performance of the GPU-accelerated eigensolver developed in this work for routine semi-local KS-DFT calculations comprising thousands of atoms.
Recommendations
- GPU acceleration for Hermitian eigensystems
- GPU optimization of large-scale eigenvalue solver
- Toward an Efficient Parallel Eigensolver for Dense Symmetric Matrices
- GPU acceleration of splitting schemes applied to differential matrix equations
- GPU acceleration of dense matrix and block operations for Lanczos method for systems over \(\mathrm{GF}(2)\)
- High-performance solvers for dense Hermitian eigenproblems
- Optimizing the computation of eigenvalues using graphics processing units
Cites work
- A Divide and Conquer method for the symmetric tridiagonal eigenproblem
- A Divide-and-Conquer Algorithm for the Symmetric Tridiagonal Eigenproblem
- A Jacobi–Davidson Iteration Method for Linear Eigenvalue Problems
- A Parallel Algorithm for Reducing Symmetric Banded Matrices to Tridiagonal Form
- A Parallel Divide and Conquer Algorithm for the Symmetric Eigenvalue Problem on Distributed Memory Architectures
- Ab initio molecular simulations with numeric atom-centered orbitals
- Accelerating numerical dense linear algebra calculations with GPUs
- Elemental, a new framework for distributed memory dense matrix computations
- ELSI: a unified software interface for Kohn-Sham electronic structure solvers
- scientific article; zbMATH DE number 6159604 (Why is no real title available?)
- Integrating state of the art compute, communication, and autotuning strategies to multiply the performance of ab initio molecular dynamics on massively parallel multi-core supercomputers
- LAPACK Users' Guide
- Massively parallel sparse matrix function calculations with NTPoly
- Numerical recipes. The art of scientific computing.
- ScaLAPACK Users' Guide
- The iterative calculation of a few of the lowest eigenvalues and corresponding eigenvectors of large real-symmetric matrices
- Towards dense linear algebra for hybrid GPU accelerated manycore systems
- Unitary Triangularization of a Nonsymmetric Matrix
Cited in
(2)
This page was built for publication: GPU-acceleration of the ELPA2 distributed eigensolver for dense symmetric and Hermitian eigenproblems
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6159210)