Regularization via mass transportation
From MaRDI portal
Abstract: The goal of regression and classification methods in supervised learning is to minimize the empirical risk, that is, the expectation of some loss function quantifying the prediction error under the empirical distribution. When facing scarce training data, overfitting is typically mitigated by adding regularization terms to the objective that penalize hypothesis complexity. In this paper we introduce new regularization techniques using ideas from distributionally robust optimization, and we give new probabilistic interpretations to existing techniques. Specifically, we propose to minimize the worst-case expected loss, where the worst case is taken over the ball of all (continuous or discrete) distributions that have a bounded transportation distance from the (discrete) empirical distribution. By choosing the radius of this ball judiciously, we can guarantee that the worst-case expected loss provides an upper confidence bound on the loss on test data, thus offering new generalization bounds. We prove that the resulting regularized learning problems are tractable and can be tractably kernelized for many popular loss functions. We validate our theoretical out-of-sample guarantees through simulated and empirical experiments.
Recommendations
- Regularization for Wasserstein distributionally robust optimization
- Robust Wasserstein profile inference and applications to machine learning
- Variance-based regularization with convex objectives
- Robust kernel-based distribution regression
- A robust learning approach for regression models based on distributionally robust optimization
Cites work
- 10.1162/153244303321897690
- 10.1162/153244303321897726
- A comment on ``Computational complexity of stochastic programming problems
- A Singular Value Thresholding Algorithm for Matrix Completion
- Applied logistic regression
- Characterization of the equivalence of robustification and regularization in linear and matrix regression
- Concentration inequalities. A nonasymptotic theory of independence
- Convex optimization theory.
- Data-driven distributionally robust optimization using the Wasserstein metric: performance guarantees and tractable reformulations
- Decomposition algorithm for distributionally robust optimization using Wasserstein metric with an application to a class of regression models
- Distributionally Robust Convex Optimization
- Distributionally robust optimization and its tractable approximations
- Distributionally robust optimization under moment uncertainty with application to data-driven problems
- Estimating the support of a high-dimensional distribution
- Finding structure with randomness: probabilistic algorithms for constructing approximate matrix decompositions
- Generalized Lagrange Multiplier Method for Solving Problems of Optimum Allocation of Resources
- scientific article; zbMATH DE number 6378089 (Why is no real title available?)
- scientific article; zbMATH DE number 3551792 (Why is no real title available?)
- scientific article; zbMATH DE number 6982911 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- Iterative Bregman projections for regularized transportation problems
- Maximum relative margin and data-dependent regularization
- On distributionally robust chance-constrained linear programs
- On duality theory of conic linear problems.
- On the rate of convergence in Wasserstein distance of the empirical measure
- Optimal Transport
- Quantifying distributional model risk via optimal transport
- Quantile regression.
- Regularisation of neural networks by enforcing Lipschitz continuity
- Robust optimization
- Robust Regression and Lasso
- Robust Solutions to Least-Squares Problems with Uncertain Data
- Robust Wasserstein profile inference and applications to machine learning
- Robustness and regularization of support vector machines
- Second order cone programming approaches for handling missing and uncertain data
- Second order cone programming formulations for feature selection
- Second-order stochastic optimization for machine learning in linear time
- Sharp asymptotic and finite-sample rates of convergence of empirical measures in Wasserstein distance
- Support-vector networks
- The earth mover's distance as a metric for image retrieval
- The minimum error minimax probability machine
- Understanding machine learning. From theory to algorithms
Cited in
(41)- Conditional variance penalties and domain shift robustness
- Superquantiles at work: machine learning applications and efficient subgradient computation
- Distributionally robust optimization. A review on theory and applications
- Deep regularization and direct training of the inner layers of neural networks with kernel flows
- Robust grouped variable selection using distributionally robust optimization
- Robust linear classification from limited training data
- Frameworks and results in distributionally robust optimization
- Partition-based distributionally robust optimization via optimal transport with order cone constraints
- On the regularized risk of distributionally robust learning over deep neural networks
- Optimal learning rates for distribution regression
- Transport via mass transportation
- A distributionally robust area under curve maximization model
- Tractable reformulations of two-stage distributionally robust linear programs over the type-\(\infty\) Wasserstein ball
- Adversarial classification via distributional robustness with Wasserstein ambiguity
- scientific article; zbMATH DE number 7370573 (Why is no real title available?)
- Distributionally robust inverse covariance estimation: the Wasserstein shrinkage estimator
- Optimal transport-based distributionally robust optimization: structural properties and iterative schemes
- Robust Wasserstein profile inference and applications to machine learning
- Generalization bounds for regularized portfolio selection with market side information
- Semi-discrete optimal transport: hardness, regularization and numerical solution
- Sensitivity of Multiperiod Optimization Problems with Respect to the Adapted Wasserstein Distance
- Regularization for Wasserstein distributionally robust optimization
- Distributionally Robust Losses for Latent Covariate Mixtures
- Solution path algorithm for distributionally robust regression
- Data-driven inverse optimization with imperfect information
- A review of algorithms for distributionally robust optimization using statistical distances
- Wasserstein distributionally robust optimization and its tractable regularization formulation
- Optimal transport regularized divergences: application to adversarial robustness
- Robust approaches in portfolio optimization with stochastic dominance constraints
- Sequential decision-making under uncertainty: a robust MDPs review
- Nonsmooth nonconvex-nonconcave minimax optimization: primal-dual balancing and iteration complexity analysis
- Distributionally robust optimization and robust statistics
- Optimal transport-based distributionally robust optimization with polynomial uncertainty
- Distributionally robust Gaussian process regression and Bayesian inverse problems
- Distributionally robust optimization
- Distributionally robust chance-constrained kernel-based support vector machine
- Supervised learning with evolving tasks and performance guarantees
- On regularization schemes for data-driven optimization based on Cressie–Read divergence and CVaR
- Sensitivity analysis of Wasserstein distributionally robust optimization problems
- Efficient data-driven optimization with noisy data
- First-order conditions for optimization in the Wasserstein space
Describes a project that uses
Uses Software
This page was built for publication: Regularization via mass transportation
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5214188)