Fast convex pruning of deep neural networks
From MaRDI portal
Abstract: We develop a fast, tractable technique called Net-Trim for simplifying a trained neural network. The method is a convex post-processing module, which prunes (sparsifies) a trained network layer by layer, while preserving the internal responses. We present a comprehensive analysis of Net-Trim from both the algorithmic and sample complexity standpoints, centered on a fast, scalable convex optimization program. Our analysis includes consistency results between the initial and retrained models before and after Net-Trim application and guarantees on the number of training samples needed to discover a network that can be expressed using a certain number of nonzero terms. Specifically, if there is a set of weights that uses at most terms that can re-create the layer outputs from the layer inputs, we can find these weights from samples, where is the input size. These theoretical results are similar to those for sparse regression using the Lasso, and our analysis uses some of the same recently-developed tools (namely recent results on the concentration of measure and convex analysis). Finally, we propose an algorithmic framework based on the alternating direction method of multipliers (ADMM), which allows a fast and simple implementation of Net-Trim for network pruning and compression.
Recommendations
- Sensitivity-informed provable pruning of neural networks
- Make _1 regularization effective in training sparse CNN
- Transformed \(\ell_1\) regularization for learning sparse deep neural networks
- Efficient and sparse neural networks by pruning weights in a multiobjective learning approach
- Dynamic schedule for effective on-line connection pruning
Cites work
- A tail inequality for quadratic forms of subgaussian random vectors
- Convex Recovery of a Structured Signal from Independent Random Linear Measurements
- Deep learning
- Distributed optimization and statistical learning via the alternating direction method of multipliers
- scientific article; zbMATH DE number 6378127 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- Learning without concentration
- Living on the edge: phase transitions in convex programs with random data
- Ridge Regression: Biased Estimation for Nonorthogonal Problems
- Stable signal recovery from incomplete and inaccurate measurements
- The Generic Chaining
- Weak convergence and empirical processes. With applications to statistics
Cited in
(14)- Training thinner and deeper neural networks: jumpstart regularization
- Max-plus operators applied to filter selection and model pruning in neural networks
- PAC-Bayesian framework based drop-path method for 2D discriminative convolutional network pruning
- Efficient and sparse neural networks by pruning weights in a multiobjective learning approach
- SVD-based DNN pruning and retraining
- Structure optimization of convolutional neural networks: a survey
- Sensitivity-informed provable pruning of neural networks
- scientific article; zbMATH DE number 7626756 (Why is no real title available?)
- slimTrain---A Stochastic Approximation Method for Training Separable Deep Neural Networks
- Dynamic schedule for effective on-line connection pruning
- Pruning during training by network efficacy modeling
- Deep Neural Networks Pruning via the Structured Perspective Regularization
- Simultaneous compression and stabilization of neural networks through pruning
- Make _1 regularization effective in training sparse CNN
This page was built for publication: Fast convex pruning of deep neural networks
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5027023)