Convolution accelerator designs using fast algorithms
Summary: Convolutional neural networks (CNNs) have achieved great success in image processing. However, the heavy computational burden it imposes makes it difficult for use in embedded applications that have limited power consumption and performance. Although there are many fast convolution algorithms that can reduce the computational complexity, they increase the difficulty of practical implementation. To overcome these difficulties, this paper proposes several convolution accelerator designs using fast algorithms. The designs are based on the field programmable gate array (FPGA) and display a better balance between the digital signal processor (DSP) and the logic resource, while also requiring lower power consumption. The implementation results show that the power consumption of the accelerator design based on the Strassen-Winograd algorithm is 21.3\% less than that of conventional accelerators.
- A faster algorithm for reducing the computational complexity of convolutional neural networks
- Method for Convolutional Neural Network Hardware Implementation Based on a Residue Number System
- A decomposable Winograd method for N-D convolution acceleration in video analysis
- Accelerating CNN models for face verification with convolution theorem
- Derivation and analysis of fast bilinear algorithms for convolution
- System-on-a-chip (SoC)-based hardware acceleration for foreground and background identification
- A decomposable Winograd method for N-D convolution acceleration in video analysis
- A faster algorithm for reducing the computational complexity of convolutional neural networks
- Arithmetic unit design for neural accelerators: cost performance issues
- Accelerating CNN models for face verification with convolution theorem
- Efficient processing of deep neural networks
- FPGA design and hardware implementation of a convolutional neural network for classification of saccadic eye movements
- Method for Convolutional Neural Network Hardware Implementation Based on a Residue Number System
- Algorithm design for tensor units
This page was built for publication: Convolution accelerator designs using fast algorithms
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2004889)