Reliable clustering of Bernoulli mixture models
From MaRDI portal
Abstract: A Bernoulli Mixture Model (BMM) is a finite mixture of random binary vectors with independent dimensions. The problem of clustering BMM data arises in a variety of real-world applications, ranging from population genetics to activity analysis in social networks. In this paper, we analyze the clusterability of BMMs from a theoretical perspective, when the number of clusters is unknown. In particular, we stipulate a set of conditions on the sample complexity and dimension of the model in order to guarantee the Probably Approximately Correct (PAC)-clusterability of a dataset. To the best of our knowledge, these findings are the first non-asymptotic bounds on the sample complexity of learning or clustering BMMs.
Recommendations
- scientific article; zbMATH DE number 194758
- Block clustering with Bernoulli mixture models: comparison of different approaches
- Finite mixture models and model-based clustering
- Clusterability assessment for Gaussian mixture models
- scientific article; zbMATH DE number 5059940
- Block Bernoulli Parsimonious Clustering Models
- A parametric mixture model for clustering multivariate binary data
- Bayesian Mixture Labeling and Clustering
- Latent mixture modeling for clustered data
Cites work
- A tutorial on Bayesian nonparametric models
- An entropy criterion for assessing the number of clusters in a mixture model
- An improvement of the NEC criterion for assessing the number of clusters in a mixture model
- Bayesian nonparametric data analysis
- Differential and integral calculus. Volume I. Transl. From the German by E. J. McShane. With a new foreword and revised by Marvin Jay Greenberg
- Efficient density estimation via piecewise polynomial approximation
- Foundations of machine learning
- scientific article; zbMATH DE number 107482 (Why is no real title available?)
- scientific article; zbMATH DE number 1222288 (Why is no real title available?)
- Identifiability of parameters in latent structure models with many observed variables
- Information Theoretical Analysis of Multivariate Correlation
- Model-based clustering
- Model-based clustering of high-dimensional data: a review
- Model-Based Clustering, Discriminant Analysis, and Density Estimation
- Near-optimal Sample Complexity Bounds for Robust Learning of Gaussian Mixtures via Compression Schemes
- Non-uniqueness in probabilistic numerical identification of bacteria
- Nonparametric statistical methods
- Pattern recognition and machine learning.
- Statistical guarantees for the EM algorithm: from population to sample-based analysis
- Structural, Syntactic, and Statistical Pattern Recognition
Cited in
(3)
This page was built for publication: Reliable clustering of Bernoulli mixture models
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2295043)