MAGMA
From MaRDI portal
MAGMA Q24666
Cited in
(only showing first 100 items - show all)- FLAME
- PaStiX
- H2Opus
- SGEMM
- ScaLAPACK
- VOLSCAT
- LAWRA
- PLAPACK
- Algorithm 826
- H-LU factorization on many-core systems
- Redesigning triangular dense matrix computations on GPUs
- Hybrid algorithms for solving the algebraic eigenvalue problem with sparse matrices
- Experiments with sparse Cholesky using a sequential task-flow implementation
- Efficient determination of the Markovian time-evolution towards a steady-state of a complex open quantum system
- CALU
- RScaLAPACK
- CUBLAS
- MKL
- POOCLAPACK
- Elemental
- OpenCL
- SOLAR
- Cellss
- SBR Toolbox
- PLASMA
- The parallel tiled WZ factorization algorithm for multicore architectures
- LU factorization on heterogeneous systems: an energy-efficient approach towards high performance
- Numerical analysis of parallel implementation of the reorthogonalized ABS methods
- clSpMV
- PLASMA
- NaSt3DGPF
- CULA
- Algorithm 880
- LogGOPSim
- MR3-SMP
- HSL_MA87
- Solving a large scale radiosity problem on GPU-based parallel computers
- Interoperable executive library for the simulation of biomedical processes
- STREAM benchmark
- IEL
- HSL_MA79
- SpGEMM
- QUARK
- StarPU
- High-order finite-element seismic wave propagation modeling with MPI on a large GPU cluster
- H2Opus: a distributed-memory multi-GPU software package for non-local operators
- Wool
- SWARM
- Scientific computations on multi-core systems using different programming frameworks
- FastFlow
- MPIGMP
- GPUprec
- CUMP
- Evaluation of selected resource allocation and scheduling methods in heterogeneous many-core processors and graphics processing units
- MINMOD
- BLIS: a framework for rapidly instantiating BLAS functionality
- Algorithm 953: Parallel library software for the multishift QR algorithm with aggressive early deflation
- An efficient multicore implementation of a novel HSS-structured multifrontal solver using randomized sampling
- ViennaCL-linear algebra library for multi- and many-core architectures
- A distributed and incremental SVD algorithm for agglomerative data analysis on large networks
- A parallel auxiliary grid algebraic multigrid method for graphic processing units
- A parallel algorithm for calculation of determinants and minors using arbitrary precision arithmetic
- Programming the finite element method
- An inertia-free filter line-search algorithm for large-scale nonlinear programming
- Divide and conquer on hybrid GPU-accelerated multicore systems
- Algorithm 953
- KBLAS
- yaSpMV
- Tcmalloc
- Exploiting symmetry in tensors for high performance: multiplication with symmetric tensors
- An efficient approach to solve very large dense linear systems with verified computing on clusters.
- DDSCAT
- Accelerating GPU kernels for dense linear algebra
- A Scalable High Performant Cholesky Factorization for Multicore with GPU Accelerators
- Sparse matrix-vector multiplication on GPGPUs
- PBLAS
- SuperMatrix
- CLBlast
- CLTune
- Algorithm 656
- DAGuE
- NFM-DS
- Computing least squares condition numbers on hybrid multicore/GPU systems
- Solving a large-scale thermal radiation problem using an interoperable executive library framework on petascale supercomputers
- Extending the length and time scales of Gram-Schmidt Lyapunov vector computations
- moderngpu
- CSR5
- BiELL
- CoAdELL
- AdELL
- PDHSEQR
- PDLAQR1
- UHM
- Multi-GPU implementation of the lattice Boltzmann method
- gptk
- GFortran
- KSVD
- Implementing High-performance Complex Matrix Multiplication via the 3m and 4m Methods
- Accelerating the solution of linear systems by iterative refinement in three precisions
- Zippy
This page was built for software: MAGMA