Data augmentation for Bayesian deep learning
From MaRDI portal
Abstract: Deep Learning (DL) methods have emerged as one of the most powerful tools for functional approximation and prediction. While the representation properties of DL have been well studied, uncertainty quantification remains challenging and largely unexplored. Data augmentation techniques are a natural approach to provide uncertainty quantification and to incorporate stochastic Monte Carlo search into stochastic gradient descent (SGD) methods. The purpose of our paper is to show that training DL architectures with data augmentation leads to efficiency gains. We use the theory of scale mixtures of normals to derive data augmentation strategies for deep learning. This allows variants of the expectation-maximization and MCMC algorithms to be brought to bear on these high dimensional nonlinear deep learning models. To demonstrate our methodology, we develop data augmentation algorithms for a variety of commonly used activation functions: logit, ReLU, leaky ReLU and SVM. Our methodology is compared to traditional stochastic gradient descent with back-propagation. Our optimization procedure leads to a version of iteratively re-weighted least squares and can be implemented at scale with accelerated linear algebra methods providing substantial improvement in speed. We illustrate our methodology on a number of standard datasets. Finally, we conclude with directions for future research.
Cites work
- scientific article; zbMATH DE number 3885116 (Why is no real title available?)
- scientific article; zbMATH DE number 3147701 (Why is no real title available?)
- scientific article; zbMATH DE number 7008320 (Why is no real title available?)
- scientific article; zbMATH DE number 6860839 (Why is no real title available?)
- scientific article; zbMATH DE number 849932 (Why is no real title available?)
- scientific article; zbMATH DE number 849933 (Why is no real title available?)
- scientific article; zbMATH DE number 3234662 (Why is no real title available?)
- A selective overview of deep learning
- Bayesian Classification of Tumours by Using Gene Expression Data
- Bayesian Deep Net GLM and GLMM
- Bayesian Inference for Logistic Models Using Pólya–Gamma Latent Variables
- Bayesian treed Gaussian process models with an application to computer modeling
- Computer Model Calibration Using High-Dimensional Output
- Data augmentation for non-Gaussian regression models using variance-mean mixtures
- Data augmentation for support vector machines
- Deep learning: a Bayesian perspective
- Deep vs. shallow networks: an approximation theory perspective
- Diffusions for Global Optimization
- Equation of state calculations by fast computing machines
- Error bounds for approximations with deep ReLU networks
- Generalized double Pareto shrinkage
- Letter to the Editor—-A Closed Form Solution of Certain Programming Problems
- Letter to the Editor—A Monte Carlo Method for the Approximate Solution of Certain Types of Constrained Optimization Problems
- MCMC maximum likelihood for latent state models
- Multivariate adaptive regression splines
- Nonparametric regression using deep neural networks with ReLU activation function
- On deep learning as a remedy for the curse of dimensionality in nonparametric regression
- Optimal Scaling of Discrete Approximations to Langevin Diffusions
- Optimization
- Sampling can be faster than optimization
- Slice sampling. (With discussions and rejoinder)
- Weighted Bayesian bootstrap for scalable posterior distributions
This page was built for publication: Data augmentation for Bayesian deep learning
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6122055)