GPipe
From MaRDI portal
Cited in
(18)- BiT
- REALM
- EGC: entropy-based gradient compression for distributed deep learning
- Binary quantized network training with sharpness-aware minimization
- A statistician teaches deep learning
- On the convergence analysis of asynchronous SGD for solving consistent linear systems
- Megatron-LM
- scientific article; zbMATH DE number 7370624 (Why is no real title available?)
- Associated learning: decomposing end-to-end backpropagation based on autoencoders and target propagation
- Deep double descent: where bigger models and more data hurt*
- The stochastic delta rule: faster and more accurate deep learning through adaptive weight noise
- GhostNet
- M2M-100
- GShard
- Mesh TensorFlow
- Europarl
- DeepSpeed
- mT5
This page was built for software: GPipe