Layer-Parallel Training of Deep Residual Neural Networks
From MaRDI portal
Recommendations
- Globally Convergent Multilevel Training of Deep Residual Networks
- Block layer decomposition schemes for training deep neural networks
- A framework for parallel and distributed training of neural networks
- Stochastic training of residual networks in deep learning
- Multilevel minimization for deep residual networks
- Parallel orthogonal deep neural network
- Deep limits of residual neural networks
Cites work
- 50 years of time parallel time integration
- A Multigrid Tutorial, Second Edition
- A non-intrusive parallel-in-time adjoint solver with the xbraid library
- A non-intrusive parallel-in-time approach for simultaneous optimization with unsteady PDEs
- A nonlinear ParaExp algorithm
- A proposal on machine learning via dynamical systems
- Adaptive multilevel inexact SQP methods for PDE-constrained optimization
- Adaptive sequencing of primal, dual, and design steps in simulation based optimization
- An Efficient Parallel-in-Time Method for Optimization with Parabolic PDEs
- An introduction to the adjoint approach to design
- Analysis of the Parareal Time‐Parallel Time‐Integration Method
- Approximate nullspace iterations for KKT systems
- Computational optimization of systems governed by partial differential equations
- Deep learning
- Evaluating Derivatives
- scientific article; zbMATH DE number 1561761 (Why is no real title available?)
- scientific article; zbMATH DE number 5060482 (Why is no real title available?)
- scientific article; zbMATH DE number 5785783 (Why is no real title available?)
- Learning deep architectures for AI
- Minimal Repetition Dynamic Checkpointing Algorithm for Unsteady Adjoint Calculation
- Multi-Level Adaptive Solutions to Boundary-Value Problems
- Multigrid methods with space-time concurrency
- Multigrid Reduction in Time for Nonlinear Parabolic Problems: A Case Study
- Natural language processing (almost) from scratch
- One-shot approaches to design optimzation
- Optimal design with bounded retardation for problems with non-separable adjoints
- Parallel Lagrange--Newton--Krylov--Schur Methods for PDE-Constrained Optimization. Part I: The Krylov--Schur Solver
- Parallel time integration with multigrid
- Stable architectures for deep neural networks
- Two-level convergence theory for multigrid reduction in time (MGRIT)
Cited in
(35)- Residual networks as flows of diffeomorphisms
- Application of the residue number system to reduce hardware costs of the convolutional neural network implementation
- Derivation and analysis of parallel-in-time neural ordinary differential equations
- Multigrid reduction in time with Richardson extrapolation
- A sequential quadratic Hamiltonian algorithm for training explicit RK neural networks
- Quantized convolutional neural networks through the lens of partial differential equations
- AutoMat: automatic differentiation for generalized standard materials on GPUs
- Decomposition and composition of deep convolutional neural networks and training acceleration via sub-network transfer learning
- Long-time integration of parametric evolution equations with physics-informed DeepONets
- Structure-preserving deep learning
- Stochastic training of residual networks in deep learning
- New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference
- A Unified Analysis Framework for Iterative Parallel-in-Time Algorithms
- Multilevel Objective-Function-Free Optimization with an Application to Neural Networks Training
- Multilevel minimization for deep residual networks
- Parallel orthogonal deep neural network
- Semi-implicit back propagation
- Globally Convergent Multilevel Training of Deep Residual Networks
- MGIC: Multigrid-in-Channels Neural Network Architectures
- Connections between numerical algorithms for PDEs and neural networks
- Efficient multigrid reduction-in-time for method-of-lines discretizations of linear advection
- Parareal with a learned coarse model for robotic manipulation
- Applications of time parallelization
- A space-time parallel algorithm with adaptive mesh refinement for computational fluid dynamics
- An optimal control framework for adaptive neural ODEs
- One-shot learning of surrogates in PDE-constrained optimization under uncertainty
- Event-based automatic differentiation of OpenMP with opDiLib
- Enhancing training of physics-informed neural networks using domain decomposition-based preconditioning strategies
- A parareal architecture for very deep convolutional neural networks
- Predict globally, correct locally: parallel-in-time optimization of neural networks
- TorchBraid: high-performance layer-parallel training of deep neural networks with MPI and GPU acceleration
- A nonoverlapping domain decomposition method for extreme learning machines: elliptic problems
- Layer-parallel training of residual networks with auxiliary variable networks
- Physics-informed gated networks for solving forward and inverse problems of evolution equations on irregular domains
- Parallel-in-time solution of Allen-Cahn equations by integrating operator learning into the parareal method
This page was built for publication: Layer-Parallel Training of Deep Residual Neural Networks
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5027015)