Nonexchangeable random partition models for microclustering
From MaRDI portal
Publication:2054470
DOI10.1214/20-AOS2003zbMATH Open1486.62066arXiv1711.07287OpenAlexW3132862470MaRDI QIDQ2054470FDOQ2054470
Authors: Giuseppe Di Benedetto, Francois Caron, Yee Whye Teh
Publication date: 3 December 2021
Published in: The Annals of Statistics (Search for Journal in Brave)
Abstract: Many popular random partition models, such as the Chinese restaurant process and its two-parameter extension, fall in the class of exchangeable random partitions, and have found wide applicability in model-based clustering, population genetics, ecology or network analysis. While the exchangeability assumption is sensible in many cases, it has some strong implications. In particular, Kingman's representation theorem implies that the size of the clusters necessarily grows linearly with the sample size; this feature may be undesirable for some applications, as recently pointed out by Miller et al. (2015). We present here a flexible class of non-exchangeable random partition models which are able to generate partitions whose cluster sizes grow sublinearly with the sample size, and where the growth rate is controlled by one parameter. Along with this result, we provide the asymptotic behaviour of the number of clusters of a given size, and show that the model can exhibit a power-law behavior, controlled by another parameter. The construction is based on completely random measures and a Poisson embedding of the random partition, and inference is performed using a Sequential Monte Carlo algorithm. Additionally, we show how the model can also be directly used to generate sparse multigraphs with power-law degree distributions and degree sequences with sublinear growth. Finally, experiments on real datasets emphasize the usefulness of the approach compared to a two-parameter Chinese restaurant process.
Full work available at URL: https://arxiv.org/abs/1711.07287
Recommendations
- Random Partition Models for Microclustering Tasks
- Clustering via nonsymmetric partition distributions
- Exchangeable and partially exchangeable random partitions
- Random partition models and exchangeability for Bayesian identification of population struc\-ture
- Random partition models with regression on covariates
Cites Work
- Sequential Monte Carlo Methods in Practice
- Sequential Monte Carlo Samplers
- Distance dependent chinese restaurant processes
- Order-Based Dependent Dirichlet Processes
- Completely random measures
- Bayesian Nonparametric Estimation of the Probability of Discovering New Species
- Particle Markov Chain Monte Carlo Methods
- Title not available (Why is that?)
- Title not available (Why is that?)
- Hierarchical Mixture Modeling With Normalized Inverse-Gaussian Priors
- The two-parameter Poisson-Dirichlet distribution derived from a stable subordinator
- Exchangeable and partially exchangeable random partitions
- Combinatorial stochastic processes. Ecole d'Eté de Probabilités de Saint-Flour XXXII -- 2002.
- Title not available (Why is that?)
- Heavy-Tail Phenomena
- Size-biased sampling of Poisson point processes and excursions
- Distributional results for means of normalized random measures with independent increments
- Normalized random measures driven by increasing additive processes
- Posterior Analysis for Normalized Random Measures with Independent Increments
- Title not available (Why is that?)
- Generalized Gamma measures and shot-noise Cox processes
- Title not available (Why is that?)
- Title not available (Why is that?)
- An Introduction to the Theory of Point Processes
- Beta processes, stick-breaking and power laws
- Title not available (Why is that?)
- Generalized Pólya urn for time-varying Pitman-Yor processes
- Contributions to the theory of Dirichlet processes
- Bayesian inference with dependent normalized completely random measures
- Title not available (Why is that?)
- Controlling the Reinforcement in Bayesian Non-Parametric Mixture Models
- Title not available (Why is that?)
- Notes on the occupancy problem with infinitely many boxes: general asymptotics and power laws
- Conditional formulae for Gibbs-type exchangeable random partitions
- Random partitions in population genetics
- Random permutations with cycle weights
- Bayesian Poisson process partition calculus with an application to Bayesian Lévy moving averages
- Looking-backward probabilities for Gibbs-type exchangeable random partitions
- Distribution theory for hierarchical processes
- Bayesian nonparametric analysis for a generalized Dirichlet process prior
- Sparse graphs using exchangeable random measures
- Nonparametric Bayesian inference.
- On edge exchangeable random graphs
- Edge exchangeable models for interaction networks
- Latent nested nonparametric priors (with discussion)
- Power laws for family sizes in a duplication model
Cited In (9)
- Clustering via nonsymmetric partition distributions
- Random partition models and complementary clustering of Anglo-Saxon place-names
- An EPPF from independent sequences of geometric random variables
- Asymptotic analysis of statistical estimators related to multigraphex processes under misspecification
- Fast Generation of Exchangeable Sequence of Clusters Data
- Asymptotic behavior of the number of distinct values in a sample from the geometric stick-breaking process
- Interfaces for random cluster models
- Concentration in the generalized Chinese restaurant process
- Random Partition Models for Microclustering Tasks
Uses Software
This page was built for publication: Nonexchangeable random partition models for microclustering
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2054470)