Measuring Complexity of Learning Schemes Using Hessian-Schatten Total Variation
From MaRDI portal
Abstract: In this paper, we introduce the Hessian-Schatten total variation (HTV) -- a novel seminorm that quantifies the total "rugosity" of multivariate functions. Our motivation for defining HTV is to assess the complexity of supervised-learning schemes. We start by specifying the adequate matrix-valued Banach spaces that are equipped with suitable classes of mixed norms. We then show that the HTV is invariant to rotations, scalings, and translations. Additionally, its minimum value is achieved for linear mappings, which supports the common intuition that linear regression is the least complex learning model. Next, we present closed-form expressions of the HTV for two general classes of functions. The first one is the class of Sobolev functions with a certain degree of regularity, for which we show that the HTV coincides with the Hessian-Schatten seminorm that is sometimes used as a regularizer for image reconstruction. The second one is the class of continuous and piecewise-linear (CPWL) functions. In this case, we show that the HTV reflects the total change in slopes between linear regions that have a common facet. Hence, it can be viewed as a convex relaxation (l1-type) of the number of linear regions (l0-type) of CPWL mappings. Finally, we illustrate the use of our proposed seminorm.
Recommendations
Cites work
- A distribution-free theory of nonparametric regression
- A second-order model for image denoising
- An introduction to sparse stochastic processes
- Approximation by superpositions of a sigmoidal function
- Banach space representer theorems for neural networks and ridge splines
- Benign overfitting in linear regression
- Compressed sensing
- Convex optimization in sums of Banach spaces
- Deep Convolutional Neural Network for Inverse Problems in Imaging
- Deep learning
- Deep learning: a statistical viewpoint
- Duality Mapping for Schatten Matrix Norms
- Fonctions à hessien borné
- Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
- Hessian Schatten-Norm Regularization for Linear Inverse Problems
- Hessian-Based Norm Regularization for Image Restoration With Biomedical Applications
- Higher-order total variation approaches and generalisations
- scientific article; zbMATH DE number 1804115 (Why is no real title available?)
- scientific article; zbMATH DE number 45848 (Why is no real title available?)
- scientific article; zbMATH DE number 1324223 (Why is no real title available?)
- scientific article; zbMATH DE number 1448982 (Why is no real title available?)
- scientific article; zbMATH DE number 967931 (Why is no real title available?)
- Impulse functions over curves and surfaces and their applications to diffraction
- Networks and the best approximation property
- Optimal rates for the regularized least-squares algorithm
- Regularization algorithms for learning that are equivalent to multilayer networks
- Regularization in kernel learning
- Regularization networks and support vector machines
- Regularization of linear inverse problems with total generalized variation
- Riesz potentials, higher Riesz transforms and Beppo Levi spaces
- Some results on Tchebycheffian spline functions and stochastic processes
- Sparsest piecewise-linear regression of one-dimensional data
- Structure tensor total variation
- Support Vector Machines
- Théorie des distributions à valeurs vectorielles
- Total generalized variation
- Variational methods on the space of functions of bounded Hessian for convexification and denoising
- What Kinds of Functions Do Deep Neural Networks Learn? Insights from Variational Spline Theory
Cited in
(3)
This page was built for publication: Measuring Complexity of Learning Schemes Using Hessian-Schatten Total Variation
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6171683)