On the regularizing property of stochastic gradient descent
From MaRDI portal
Abstract: Stochastic gradient descent is one of the most successful approaches for solving large-scale problems, especially in machine learning and statistics. At each iteration, it employs an unbiased estimator of the full gradient computed from one single randomly selected data point. Hence, it scales well with problem size and is very attractive for truly massive dataset, and holds significant potentials for solving large-scale inverse problems. In the recent literature of machine learning, it was empirically observed that when equipped with early stopping, it has regularizing property. In this work, we rigorously establish its regularizing property (under extit{a priori} early stopping rule), and also prove convergence rates under the canonical sourcewise condition, for minimizing the quadratic functional for linear inverse problems. This is achieved by combining tools from classical regularization theory and stochastic analysis. Further, we analyze the preasymptotic weak and strong convergence behavior of the algorithm. The theoretical findings shed insights into the performance of the algorithm, and are complemented with illustrative numerical experiments.
Recommendations
- On the Convergence of Stochastic Gradient Descent for Linear Inverse Problems in Banach Spaces
- On the Convergence of Stochastic Gradient Descent for Nonlinear Ill-Posed Problems
- Stochastic gradient descent for linear inverse problems in Hilbert spaces
- A Markov Chain Theory Approach to Characterizing the Minimax Optimality of Stochastic Gradient Descent (for Least Squares)
- On the regularization effect of stochastic gradient descent applied to least-squares
Cites work
- A randomized Kaczmarz algorithm with exponential convergence
- A Stochastic Approximation Method
- Acceleration of Stochastic Approximation by Averaging
- Consistency and rates of convergence of nonlinear Tikhonov regularization with random noise
- Convergence Rates of General Regularization Methods for Statistical Inverse Problems and Applications
- Heuristic Parameter-Choice Rules for Convex Variational Regularization Based on Error Estimates
- scientific article; zbMATH DE number 936298 (Why is no real title available?)
- Inverse problems. Tikhonov theory and algorithms
- Iterative regularization methods for nonlinear ill-posed problems
- Iteratively Regularized Gauss–Newton Method for Nonlinear Inverse Problems with Random Noise
- Nonparametric stochastic approximation with large step-sizes
- On the Adaptive Selection of the Parameter in Regularization of Ill-Posed Problems
- Online gradient descent learning algorithms
- Online Learning as Stochastic Approximation of Regularization Paths: Optimality and Almost-Sure Convergence
- Online learning in optical tomography: a stochastic approach
- Optimal rates for multi-pass stochastic gradient methods
- Optimization methods for large-scale machine learning
- Preasymptotic convergence of randomized Kaczmarz method
- Robust Stochastic Approximation Approach to Stochastic Programming
- Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
- The mathematics of computerized tomography
Cited in
(36)- On the regularization effect of stochastic gradient descent applied to least-squares
- Optimal rates for multi-pass stochastic gradient methods
- Convergence analyses based on frequency decomposition for the randomized row iterative method
- Randomized Kaczmarz Converges Along Small Singular Vectors
- Making the last iterate of SGD information theoretically optimal
- Online stochastic gradient descent on non-convex losses from high-dimensional inference
- An analysis of stochastic variance reduced gradient for linear inverse problems *
- Two-Layer Neural Networks with Values in a Banach Space
- Stochastic asymptotical regularization for linear inverse problems
- The inverse variance-flatness relation in stochastic gradient descent is critical for finding flat minima
- Stochastic gradient descent for linear inverse problems in Hilbert spaces
- On the Convergence of Stochastic Gradient Descent for Nonlinear Ill-Posed Problems
- On the discrepancy principle for stochastic gradient descent
- Kalman-based stochastic gradient method with stop condition and insensitivity to conditioning
- Implicit regularization with strongly convex bias: Stability and acceleration
- Stochastic mirror descent method for linear ill-posed problems in Banach spaces
- Stochastic linear regularization methods: random discrepancy principle and applications
- On the Convergence of Stochastic Gradient Descent for Linear Inverse Problems in Banach Spaces
- Max-affine regression via first-order methods
- Randomized progressive iterative approximation for B-spline curve and surface fittings
- Greedy randomized Kaczmarz with momentum method for nonlinear equation
- Stochastic asymptotical regularization for nonlinear ill-posed problems
- Convergence analysis of a stochastic heavy-ball method for linear ill-posed problems
- Stochastic variance reduced gradient method for linear ill-posed inverse problems
- A guide to stochastic optimisation for large-scale inverse problems
- Stochastic gradient descent method with convex penalty for ill-posed problems in Banach spaces
- Stochastic data-driven Bouligand-Landweber method for solving non-smooth inverse problems
- On early stopping of stochastic mirror descent method for ill-posed inverse problems
- On the convergence of a data-driven regularized stochastic gradient descent for nonlinear ill-posed problems
- On projective stochastic-gradient type methods for solving large scale systems of nonlinear ill-posed equations: applications to machine learning
- On the convergence of stochastic variance reduced gradient for linear inverse problems
- Randomized Krylov-Projected Iterated Tikhonov Regularization for Large-Scale Ill-posed Problems Under A Posteriori Stopping Rule
- Convergence analysis of the ADAM algorithm for linear inverse problems
- Randomization in inverse problems in Hilbert spaces
- Stochastic gradient descent for nonlinear inverse problems in Banach spaces
- Early stopping of stochastic variance reduced gradient for linear inverse problems by the discrepancy principle
This page was built for publication: On the regularizing property of stochastic gradient descent
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4646419)