A simple approach for quantizing neural networks
From MaRDI portal
Abstract: In this short note, we propose a new method for quantizing the weights of a fully trained neural network. A simple deterministic pre-processing step allows us to quantize network layers via memoryless scalar quantization while preserving the network performance on given training data. On one hand, the computational complexity of this pre-processing slightly exceeds that of state-of-the-art algorithms in the literature. On the other hand, our approach does not require any hyper-parameter tuning and, in contrast to previous methods, allows a plain analysis. We provide rigorous theoretical guarantees in the case of quantizing single network layers and show that the relative error decays with the number of parameters in the network if the training data behaves well, e.g., if it is sampled from suitable random distributions. The developed method also readily allows the quantization of deep networks by consecutive application to single layers.
Cites work
- A Deterministic Linear Program Solver in Current Matrix Multiplication Time
- A greedy algorithm for quantizing neural networks
- Approximation with one-bit polynomials in Bernstein form
- Deep learning
- High-dimensional probability. An introduction with applications in data science
- High-dimensional statistics. A non-asymptotic viewpoint
Cited in
(10)- Quantization for distributed estimation using neural networks
- Stochastic quantization for learning accurate low-bit deep neural networks
- QUANTIZED DETECTOR NETWORKS: A REVIEW OF RECENT DEVELOPMENTS
- Network Vector Quantization
- scientific article; zbMATH DE number 1342303 (Why is no real title available?)
- scientific article; zbMATH DE number 6982943 (Why is no real title available?)
- Frame quantization of neural networks
- Unified stochastic framework for neural network quantization and pruning
- Computability of classification and deep learning: from theoretical limits to practical feasibility through quantization
- SPFQ: a stochastic algorithm and its error analysis for neural network quantization
This page was built for publication: A simple approach for quantizing neural networks
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6172170)