Introducing a parallel cache oblivious blocking approach for the lattice Boltzmann method
Summary: We propose a parallel cache oblivious spatial and temporal blocking algorithm for the lattice Boltzmann method in three spatial dimensions. The algorithm has originally been proposed by \textit{M. Frigo} et al. [ACM Trans. Algorithms 8, No. 1, Paper No. 4, 22 p. (2012; Zbl 1295.68236)] and divides the space-time domain of stencil-based methods in an optimal way, independently of any external parameters, e.g., cache size. In view of the increasing gap between processor speed and memory performance this approach offers a promising path to increase cache utilisation. We find that even a straightforward cache oblivious implementation can reduce memory traffic at least by a factor of two if compared to a highly optimised standard kernel and improves scalability for shared memory parallelisation. Due to the recursive structure of the algorithm we use an unconventional parallelisation scheme based on task queuing.
- Parallel lattice Boltzmann method with blocked partitioning
- A parallel workload balanced and memory efficient lattice-Boltzmann algorithm with single unit BGK relaxation time for laminar Newtonian flows
- Towards a hybrid parallelization of lattice Boltzmann methods
- Auto-vectorization friendly parallel lattice Boltzmann streaming scheme for direct addressing
- scientific article; zbMATH DE number 2009920
- Parallel lattice Boltzmann method with blocked partitioning
- OpenLB -- open source lattice Boltzmann code
- \textsc{waLBerla}: a block-structured high-performance framework for multiphysics simulations
- scientific article; zbMATH DE number 2009920 (Why is no real title available?)
- Multicore-optimized wavefront diamond blocking for optimizing stencil updates
- Heterogeneous LBM Simulation Code with LRnLA Algorithms
- Designing a 3D parallel memory-aware lattice Boltzmann algorithm on manycore systems
- Towards a hybrid parallelization of lattice Boltzmann methods
This page was built for publication: Introducing a parallel cache oblivious blocking approach for the lattice Boltzmann method
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q929255)