CUBLAS
From MaRDI portal
Cited in
(only showing first 100 items - show all)- ITPACK
- Ensign
- KronPACK
- SGEMM
- VOLSCAT
- LAWRA
- BLAS
- CUDA
- Discrete particle swarm optimization for constructing uniform design on irregular regions
- MPI-CUDA sparse matrix-vector multiplication for the conjugate gradient method with an approximate inverse preconditioner
- GPU accelerated computational homogenization based on a variational approach in a reduced basis framework
- A parallel computing method using blocked format with optimal partitioning for SpMV on GPU
- GPU accelerated intensities MPI (GAIN-MPI): a new method of computing Einstein-A coefficients
- Redesigning triangular dense matrix computations on GPUs
- Efficient determination of the Markovian time-evolution towards a steady-state of a complex open quantum system
- Cucheb: a GPU implementation of the filtered Lanczos procedure
- GPU-accelerated algorithms for many-particle continuous-time quantum walks
- A new efficient and accurate spline algorithm for the matrix exponential computation
- An efficient and accurate algorithm for computing the matrix cosine based on new Hermite approximations
- GPGPU-based parallel computing applied in the FEM using the conjugate gradient algorithm: a review
- SHTns
- P3DFFT
- RScaLAPACK
- MKL
- Seigtool
- OpenCL
- Algorithm 919
- CUSP
- PFFT
- CUSPARSE
- Higher order finite elements in space and time for anisotropic simulations with variational integrators. Application of an efficient GPU implementation
- GPU optimization of large-scale eigenvalue solver
- Efficient L₀ resampling of point sets
- SciPAL
- TERMOFLUIDS
- High-performance statistical computing in the computing environments of the 2020s
- PyCUDA
- Thrust
- HPMaX: heterogeneous parallel matrix multiplication using CPUs and GPUs
- Parallel reduction of four matrices to condensed form for a generalized matrix eigenvalue algorithm
- SIMPAR
- A GPU application for high-order compact finite difference scheme
- GPU-based block-wise nonlocal means denoising for 3D ultrasound images
- 3D data denoising via nonlocal means filter by using parallel GPU strategies
- gem5
- A heterogeneous parallel LU factorization algorithm based on a basic column block uniform allocation strategy
- OpenACC
- cuFFT
- Fast and robust flow simulations in discrete fracture networks with gpgpus
- cuRAND
- Fast Taylor polynomial evaluation for the computation of the matrix cosine
- OpenBLAS
- clSpMV
- MAGMA
- CULA
- SoftFloat
- MERAM
- Rapid re-meshing and re-solution of three-dimensional boundary element problems for interactive stress analysis
- AmgX
- GAMPACK
- CORAL
- IEL
- MPC Toolbox
- SpGEMM
- PyOpenCL
- SeLaLib
- QUARK
- gputools
- CONLIN
- testmatrix
- Compressed hierarchical Schur algorithm for frequency-domain analysis of photonic structures
- CholeskyQR2
- New Hermite series expansion for computing the matrix hyperbolic cosine
- Optimal size of the block in block GMRES on GPUs: computational model and experiments
- Solving time-fractional reaction-diffusion systems through a tensor-based parallel algorithm
- A \(\mu\)-mode BLAS approach for multidimensional tensor-structured problems
- Development of a parallel CUDA algorithm for solving 3D guiding center problems
- Highly efficient GPU eigensolver for three-dimensional photonic crystal band structures with any Bravais lattice
- SCELib4.0: the new program version for computing molecular properties in the single center approach
- Boda-RTC
- AccFFT
- UPC++
- Accelerated dimension-independent adaptive metropolis
- A fast dense triangular solve in CUDA
- pyCTQW
- Sailfish
- KBLAS
- AUGEM
- Parallel and Heterogeneous m--Hessenberg--Triangular--Triangular Reduction
- cuDNN
- A Runtime System for Programming Out-of-Core Matrix Algorithms-by-Tiles on Multithreaded Architectures
- Performance and numerical accuracy evaluation of heterogeneous multicore systems for Krylov orthogonal basis computation
- An error correction solver for linear systems: evaluation of mixed precision implementations
- Accelerating GPU kernels for dense linear algebra
- A Scalable High Performant Cholesky Factorization for Multicore with GPU Accelerators
- Accelerating the explicitly restarted Arnoldi method with GPUs using an autotuned matrix vector product
- Efficient and accurate algorithms for computing matrix trigonometric functions
- Exposing fine-grained parallelism in algebraic multigrid methods
- Performance models and workload distribution algorithms for optimizing a hybrid CPU-GPU multifrontal solver
- maxDNN
This page was built for software: CUBLAS