Learning theory for distribution regression
From MaRDI portal
Abstract: We focus on the distribution regression problem: regressing to vector-valued outputs from probability measures. Many important machine learning and statistical tasks fit into this framework, including multi-instance learning and point estimation problems without analytical solution (such as hyperparameter or entropy estimation). Despite the large number of available heuristics in the literature, the inherent two-stage sampled nature of the problem makes the theoretical analysis quite challenging, since in practice only samples from sampled distributions are observable, and the estimates have to rely on similarities computed between sets of points. To the best of our knowledge, the only existing technique with consistency guarantees for distribution regression requires kernel density estimation as an intermediate step (which often performs poorly in practice), and the domain of the distributions to be compact Euclidean. In this paper, we study a simple, analytically computable, ridge regression-based alternative to distribution regression, where we embed the distributions to a reproducing kernel Hilbert space, and learn the regressor from the embeddings to the outputs. Our main contribution is to prove that this scheme is consistent in the two-stage sampled setup under mild conditions (on separable topological domains enriched with kernels): we present an exact computational-statistical efficiency trade-off analysis showing that our estimator is able to match the one-stage sampled minimax optimal rate [Caponnetto and De Vito, 2007; Steinwart et al., 2009]. This result answers a 17-year-old open question, establishing the consistency of the classical set kernel [Haussler, 1999; Gaertner et. al, 2002] in regression. We also cover consistency for more recent kernels on distributions, including those due to [Christmann and Steinwart, 2010].
Recommendations
Cited in
(37)- A class of optimal estimators for the covariance operator in reproducing kernel Hilbert spaces
- Hyperlink regression via Bregman divergence
- On a regularization of unsupervised domain adaptation in RKHS
- Minimax rate for optimal transport regression between distributions
- Distributional anchor regression
- Learning rate of distribution regression with dependent samples
- Distributed learning and distribution regression of coefficient regularization
- Optimal learning rates for distribution regression
- Multivariate tests of independence based on a new class of measures of independence in reproducing kernel Hilbert space
- Kernel regression, minimax rates and effective dimensionality: beyond the regular case
- Learning Structure Illuminates Black Boxes – An Introduction to Estimation of Distribution Algorithms
- Kernel method for persistence diagrams via kernel embedding and weight factor
- Characteristic and universal tensor product kernels
- Domain generalization by marginal transfer learning
- Marginally Calibrated Deep Distributional Regression
- Distribution regression model with a Reproducing Kernel Hilbert Space approach
- GAMLSS: A distributional regression approach
- Robust kernel-based distribution regression
- scientific article; zbMATH DE number 7415114 (Why is no real title available?)
- Computing functions of random variables via reproducing kernel Hilbert space representations
- Wasserstein Regression
- Domain Generalization by Functional Regression
- Estimates on learning rates for multi-penalty distribution regression
- Deep learning theory of distribution regression with CNNs
- Characteristic kernels on Hilbert spaces, Banach spaces, and on sets of measures
- Direct Bayesian linear regression for distribution-valued covariates
- Covariance parameter estimation of Gaussian processes with approximated functional inputs
- The impact of smoothness of kernels and target functions on unsupervised covariate shift adaptation in RKHS
- Signature methods in machine learning
- Statistical learning on measures: an application to persistence diagrams
- On the convergence rate of two-stage sampling distribution regression
- Distributional Outcome Regression via Quantile Functions and its Application to Modelling Continuously Monitored Heart Rate and Physical Activity
- Improved learning theory for kernel distribution regression with two-stage sampling
- Learning theory of distribution regression with neural networks
- Embedding distributional data
- Universal representation of permutation-invariant functions on vectors and tensors
- Deep distribution regression
This page was built for publication: Learning theory for distribution regression
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2834480)