Distance-based and RKHS-based dependence metrics in high dimension

From MaRDI portal



Abstract: In this paper, we study distance covariance, Hilbert-Schmidt covariance (aka Hilbert-Schmidt independence criterion [Gretton et al. (2008)]) and related independence tests under the high dimensional scenario. We show that the sample distance/Hilbert-Schmidt covariance between two random vectors can be approximated by the sum of squared componentwise sample cross-covariances up to an asymptotically constant factor, which indicates that the distance/Hilbert-Schmidt covariance based test can only capture linear dependence in high dimension. As a consequence, the distance correlation based t-test developed by Szekely and Rizzo (2013) for independence is shown to have trivial limiting power when the two random vectors are nonlinearly dependent but component-wisely uncorrelated. This new and surprising phenomenon, which seems to be discovered for the first time, is further confirmed in our simulation study. As a remedy, we propose tests based on an aggregation of marginal sample distance/Hilbert-Schmidt covariances and show their superior power behavior against their joint counterparts in simulations. We further extend the distance correlation based t-test to those based on Hilbert-Schmidt covariance and marginal distance/Hilbert-Schmidt covariance. A novel unified approach is developed to analyze the studentized sample distance/Hilbert-Schmidt covariance as well as the studentized sample marginal distance covariance under both null and alternative hypothesis. Our theoretical and simulation results shed light on the limitation of distance/Hilbert-Schmidt covariance when used jointly in the high dimensional setting and suggest the aggregation of marginal distance/Hilbert-Schmidt covariance as a useful alternative.


The authors consider conditions under which \[\mathrm{d}\mathrm{Cov}_n^2\left(\mathbf{X},\mathbf{Y}\right)\approx\frac{1}{\tau}\sum_{i=1}^p\sum_{j=1}^q\mathrm{cov}_n^2\left(\mathcal{X}_i,\mathcal{Y}_j\right).\] Here \(\mathrm{d}\mathrm{Cov}_n^2\left(\mathbf{X},\mathbf{Y}\right)\) denotes the umbiased sample distance covariance, \[ \mathbf{X}=\left(X_1,X_2,\ldots,X_n\right)^\intercal =\left(\mathcal{X}_1,\mathcal{X}_2,\ldots,\mathcal{X}_p\right), \] \[ \mathbf{Y}=\left(Y_1,Y_2,\ldots,Y_n\right)^\intercal=\left(\mathcal{Y}_1,\mathcal{Y}_2,\ldots,\mathcal{Y}_q\right) \] are the sample matrices, where \(X_k\mathop{=}\limits^{d}X\) and \(Y_k\mathop{=}\limits^{d}Y\) are independent samples of two random vectors \(X=(x_1,x_2,\ldots,x_p)\in\mathbb{R}^p\) and \(Y=(y_1,y_2,\ldots,y_q)\in\mathbb{R}^q\) with finite componentwise second moments, and \(\mathcal{X}_i\), \(\mathcal{Y}_j\) are the componentwise samples. In addition, \(\tau\) denotes a quantity depending on the marginal distributions of \(X\) and \(Y\) as well as \(p\) and \(q\), and \(\mathrm{cov}_n\left(\mathcal{X}_i,\mathcal{Y}_j\right)\) is an unbiased sample estimate of \(\mathrm{cov}(x_i,y_j)\). The above approximate equality is considered as \(p,q\) tend to infinity, and \(n\) can either be fixed or grows to infinity at a slower rate. The paper is the first work on the connection between sample distance covariance and sample covariance.



Cites work


Cited in
(40)








This page was built for publication: Distance-based and RKHS-based dependence metrics in high dimension

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1996774)