Learning quantized neural nets by coarse gradient method for nonlinear classification
From MaRDI portal
Publication:2050846
Abstract: Quantized or low-bit neural networks are attractive due to their inference efficiency. However, training deep neural networks with quantized activations involves minimizing a discontinuous and piecewise constant loss function. Such a loss function has zero gradients almost everywhere (a.e.), which makes the conventional gradient-based algorithms inapplicable. To this end, we study a novel class of emph{biased} first-order oracle, termed coarse gradient, for overcoming the vanished gradient issue. A coarse gradient is generated by replacing the a.e. zero derivatives of quantized (i.e., stair-case) ReLU activation composited in the chain rule with some heuristic proxy derivative called straight-through estimator (STE). Although having been widely used in training quantized networks empirically, fundamental questions like when and why the ad-hoc STE trick works, still lacks theoretical understanding. In this paper, we propose a class of STEs with certain monotonicity, and consider their applications to the training of a two-linear-layer network with quantized activation functions for non-linear multi-category classification. We establish performance guarantees for the proposed STEs by showing that the corresponding coarse gradient methods converge to the global minimum, which leads to a perfect classification. Lastly, we present experimental results on synthetic data as well as MNIST dataset to verify our theoretical findings and demonstrate the effectiveness of our proposed STEs.
Recommendations
- Blended coarse gradient descent for full quantization of deep neural networks
- Stochastic Markov gradient descent and training low-bit neural networks
- scientific article; zbMATH DE number 6982943
- BinaryRelax: a relaxation approach for training deep neural networks with quantized weights
- Stochastic quantization for learning accurate low-bit deep neural networks
Cites work
- BinaryRelax: a relaxation approach for training deep neural networks with quantized weights
- Blended coarse gradient descent for full quantization of deep neural networks
- scientific article; zbMATH DE number 6982943 (Why is no real title available?)
- scientific article; zbMATH DE number 3231758 (Why is no real title available?)
- Large margin classification using the perceptron algorithm
- Linear feature transform and enhancement of classification on deep neural network
- ReLU deep neural networks and linear finite elements
Cited in
(12)- Stochastic Markov gradient descent and training low-bit neural networks
- Recurrence of optimum for training weight and activation quantized networks
- Binary quantized network training with sharpness-aware minimization
- Stochastic quantization for learning accurate low-bit deep neural networks
- Blended coarse gradient descent for full quantization of deep neural networks
- Self-organization of the batch Kohonen network under quantization effects
- scientific article; zbMATH DE number 6982943 (Why is no real title available?)
- Learning Multiple Quantiles With Neural Networks
- How many bits does it take to quantize your neural network?
- BinaryRelax: a relaxation approach for training deep neural networks with quantized weights
- Neural Quadratic Discriminant Analysis: Nonlinear Decoding with V1-Like Computation
- scientific article; zbMATH DE number 7733439 (Why is no real title available?)
Describes a project that uses
Uses Software
This page was built for publication: Learning quantized neural nets by coarse gradient method for nonlinear classification
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2050846)