Divergence-based estimation and testing of statistical models of classification
A frequent problem of categorical data analysis is that a fixed number \(n\) of samples \(X = (X_1, \dots, X_n) \in {\mathcal X}^n\) is taken from each of \(N\) different populations (families of individuals, clusters of objects). The sample space \(\mathcal X\) is classified into \(r\) categories by a rule \(\rho : {\mathcal X} \to \{1, \dots, r\}\). Let \(Y = (Y_1,\dots, Y_r)\) be the classification vector with the components representing counts of the respective categories in the sample vector \(X\); i.e., let \[ Y_j = \#\{1 \leq k \leq n: \rho(X_k) = j\},\quad 1 \leq j\leq r. \] The sample space of the vector \(Y\) is denoted by \(S_{n,r}\); i.e., \[ S_{n,r} = \{y = (y_1,\dots, y_r) \in \{0,1,\dots, n\}^r : y_1 + \cdots + y_r = n\}. \] Populations \(i =1,\dots, N\) generate different sampel vectors \(X^{(i)}\) and the corresponding classification vectors \(Y^{(i)}\). The sampled populations are assumed to be independent and homogeneous in the sense that \(X^{(i)}\), and consequently \(Y^{(i)}\), are independent realizations of the above considered \(X\) and \(Y\). The i.i.d. property of the components \(X_1,\dots, X_n\) is included as a special case. The aim of this paper is to present an extended class of methods for estimating parameters of statistical models of vectors \(Y\) and for testing statistical hypotheses about these models. Our methods are based on the so-called \(\phi\)-divergences of probability distributions. They include as particular cases the well-known maximum likelihood method of estimation and Pearson's \(X^2\)-method of testing. Asymptotic properties of estimators minimizing \(\phi\)-divergence between theoretical and empirical vectors of means are established. Asymptotic distributions of \(\phi\)-divergences between empirical and estimated vectors of means are explicitly evaluated, and tests based on these statistics are studied.
- Minimum \(K_\phi\)-divergence estimator.
- The \(K_{\varphi}\)-divergence statistic for categorical data problems
- Divergence-based estimation and testing with misclassified data
- Phi-divergences and polytomous logistic regression models: An overview
- Statistical inference in constrained latent class models for multinomial data based on \(\phi\)-divergence measures
- A refined Jensen's inequality in Hilbert spaces and empirical approximations
- Informational distances and related statistics in mixed continuous and categorical variables
- Some new statistics for testing hypotheses in parametric models
- Asymptotic laws for disparity statistics in product multinomial models.
- New smooth test statistics of goodness-of-fit for categorized composite null hypotheses
- A generalized \(\varphi\)-divergence for asymptotically multivariate normal models.
- Some approximations to power functions of -divergence tests in parametric models
- New statistics to test log-linear modeling hypothesis with no distributional specifications and clusters with homogeneous correlation
- scientific article; zbMATH DE number 1208124 (Why is no real title available?)
- scientific article; zbMATH DE number 1234637 (Why is no real title available?)
- Asymptotic approximations for the distributions of the (h‐, φ‐)‐divergence goodness‐of‐fit statistics: application to Renyi’s statistic
- About divergence-based goodness-of-fit tests in the dirichlet-multinomial model
- scientific article; zbMATH DE number 1014737 (Why is no real title available?)
- scientific article; zbMATH DE number 1087963 (Why is no real title available?)
- Two approaches to grouping of data and related disparity statistics
- Rényi Statistics in Directed Families of Exponential Experiments*
- Empirically Estimable Classification Bounds Based on a Nonparametric Divergence Measure
- Divergence-based tests of homogeneity for spatial data
- Inference about stationary distributions of Markov chains based on divergences with observed frequencies.
- New improved estimators for overdispersion in models with clustered multinomial data and unequal cluster sizes
- Adjusted power-divergence test statistics for simple goodness-of-fit under clustered sampling
- Comments on: ``Deville and Särndal's calibration: revisiting a 25 years old successful optimization problem
This page was built for publication: Divergence-based estimation and testing of statistical models of classification
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1898412)