An experimental comparison of model-based clustering methods
From MaRDI portal
Abstract: We examine methods for clustering in high dimensions. In the first part of the paper, we perform an experimental comparison between three batch clustering algorithms: the Expectation-Maximization (EM) algorithm, a winner take all version of the EM algorithm reminiscent of the K-means algorithm, and model-based hierarchical agglomerative clustering. We learn naive-Bayes models with a hidden root node, using high-dimensional discrete-variable data sets (both real and synthetic). We find that the EM algorithm significantly outperforms the other methods, and proceed to investigate the effect of various initialization schemes on the final solution produced by the EM algorithm. The initializations that we consider are (1) parameters sampled from an uninformative prior, (2) random perturbations of the marginal distribution of the data, and (3) the output of hierarchical agglomerative clustering. Although the methods are substantially different, they lead to learned models that are strikingly similar in quality.
Recommendations
- Model-based clustering
- Model-based clustering
- 10.1162/1532443041827943
- Studying Complexity of Model-based Clustering
- Model-Based Gaussian and Non-Gaussian Clustering
- Model based clustering for mixed data: clustMD
- Model-based evaluation of clustering validation measures
- Model-based clustering of high-dimensional data: a review
Cited in
(25)- Improved model-based clustering performance using Bayesian initialization averaging
- Comparing clusterings using combination of the kappa statistic and entropy-based measure
- External clustering validity index based on chi-squared statistical test
- Finite mixtures of unimodal beta and gamma densities and the k-bumps algorithm
- Comparing clusterings -- an information based distance
- A comparative study of the K-means algorithm and the normal mixture model for clustering: univariate case
- Speed-up for the expectation-maximization algorithm for clustering categorical data
- An empirical comparison and characterisation of nine popular clustering methods
- Studying Complexity of Model-based Clustering
- Bayesian k-Means as a “Maximization-Expectation” Algorithm
- scientific article; zbMATH DE number 1977810 (Why is no real title available?)
- 10.1162/153244302760200678
- Performance evaluation of compromise conditional Gaussian networks for data clustering
- Benchmarking distance-based partitioning methods for mixed-type data
- Initialization of Hidden Markov and Semi‐Markov Models: A Critical Evaluation of Several Strategies
- A fair-multicluster approach to clustering of categorical data
- A bootstrap-based aggregate classifier for model-based clustering
- Image Comparison Based On Local Pixel Clustering
- Clustering risk in nonparametric hidden Markov and i.i.d. models
- Normalised clustering accuracy: an asymmetric external cluster validity measure
- A hypothesis test for comparing two partitions obtained from the same dataset
- Semi-supervised model-based document clustering: a comparative study
- Instability and cluster stability variance for real clusterings
- Choosing starting values for the EM algorithm for getting the highest likelihood in multivariate Gaussian mixture models
- Complexity control in a mixture model by the Hardy-Weinberg equilibrium
This page was built for publication: An experimental comparison of model-based clustering methods
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5928968)