BLIS: a framework for rapidly instantiating BLAS functionality
From MaRDI portal
Recommendations
- Analytical modeling is enough for high-performance BLIS
- scientific article; zbMATH DE number 1424342
- The BLAS API of BLASFEO: optimizing performance for small matrices
- A framework for high-performance matrix multiplication based on hierarchical abstractions, algorithms and optimized low-level kernels
- The RISC BLAS
Cites work
- scientific article; zbMATH DE number 1728263 (Why is no real title available?)
- A Storage-Efficient WY Representation for Products of Householder Transformations
- A set of level 3 basic linear algebra subprograms
- Accumulating Householder transformations, revisited
- Algorithm-Based Fault Tolerance for Matrix Operations
- An extended set of FORTRAN basic linear algebra subprograms
- An overview of the sparse basic linear algebra subprograms
- Anatomy of high-performance matrix multiplication
- Basic Linear Algebra Subprograms for Fortran Usage
- Cache efficient bidiagonalization using BLAS 2.5 operators
- Codesign Tradeoffs for High-Performance, Low-Power Linear Algebra Architectures
- Elemental, a new framework for distributed memory dense matrix computations
- Exploiting symmetry in tensors for high performance: multiplication with symmetric tensors
- FLAME
- Families of Algorithms for Reducing a Matrix to Condensed Form
- GEMM-based level 3 BLAS
- LAPACK Users' Guide
- Programming matrix algorithms-by-blocks for thread-level parallelism
- Restructuring the tridiagonal and bidiagonal QR algorithms for performance
- The WY Representation for Products of Householder Matrices
- The science of deriving dense linear algebra algorithms
- Towards an efficient tile matrix inversion of symmetric positive definite matrices on multicore architectures
Cited in
(28)- Householder QR factorization with randomization for column pivoting (HQRRP)
- Analytical modeling of matrix–vector multiplication on multicore processors
- Towards an efficient use of the BLAS library for multilinear tensor contractions
- Parallel direct solver for solving systems of linear equations resulting from finite element method on multi-core desktops and workstations
- Optimized implementation for calculation and fast-update of Pfaffians installed to the open-source fermionic variational solver mVMC
- High-Performance Tensor Contraction without Transposition
- Surrogate-based autotuning for randomized sketching algorithms in regression problems
- GMRES with embedded ensemble propagation for the efficient solution of parametric linear systems in uncertainty quantification of computational models
- A compute-bound formulation of Galerkin model reduction for linear time-invariant dynamical systems
- Efficiency of reproducible level 1 BLAS
- Parameter estimation via time modeling for MLIR implementation of GEMM
- Which C compiler and BLAS/LAPACK library should i use: \texttt{gretl}'s numerical efficiency in different configurations
- A high-performance implementation of atomistic spin dynamics simulations on x86 CPUs
- Cache optimization and performance modeling of batched, small, and rectangular matrix multiplication on Intel, AMD, and Fujitsu processors
- Algorithm 1039: automatic generators for a family of matrix multiplication routines with Apache TVM
- BLASFEO: Basic linear algebra subroutines for embedded optimization
- The BLAS API of BLASFEO: optimizing performance for small matrices
- Accelerating BLAS and LAPACK via efficient floating point architecture design
- The matrix reloaded: multiplication strategies in FrodoKEM
- BLIS
- Strassen's Algorithm for Tensor Contraction
- Implementing high-performance complex matrix multiplication via the 1M method
- Replicated computational results (RCR) report for ``BLIS: a framework for rapidly instantiating BLAS functionality
- Multidimensional Array Data Management
- Analytical modeling is enough for high-performance BLIS
- An efficient implementation of two-component relativistic density functional theory with torque-free auxiliary variables
- Extension of accurate numerical algorithms for matrix multiplication based on error-free transformation
- High-performance matrix-free unfitted finite element operator evaluation
Describes a project that uses
Uses Software
This page was built for publication: BLIS: a framework for rapidly instantiating BLAS functionality
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2828133)