Fast approximate inference for variable selection in Dirichlet process mixtures, with an application to pan-cancer proteomics
From MaRDI portal
Publication:2195280
Abstract: The Dirichlet Process (DP) mixture model has become a popular choice for model-based clustering, largely because it allows the number of clusters to be inferred. The sequential updating and greedy search (SUGS) algorithm (Wang and Dunson, 2011) was proposed as a fast method for performing approximate Bayesian inference in DP mixture models, by posing clustering as a Bayesian model selection (BMS) problem and avoiding the use of computationally costly Markov chain Monte Carlo methods. Here we consider how this approach may be extended to permit variable selection for clustering, and also demonstrate the benefits of Bayesian model averaging (BMA) in place of BMS. Through an array of simulation examples and well-studied examples from cancer transcriptomics, we show that our method performs competitively with the current state-of-the-art, while also offering computational benefits. We apply our approach to reverse-phase protein array (RPPA) data from The Cancer Genome Atlas (TCGA) in order to perform a pan-cancer proteomic characterisation of 5,157 tumour samples. We have implemented our approach, together with the original SUGS algorithm, in an open-source R package named sugsvarsel, which accelerates analysis by performing intensive computations in C++ and provides automated parallel processing. The R package is freely available from: https://github.com/ococrook/sugsvarsel
Recommendations
- A comparative review of variable selection techniques for covariate dependent Dirichlet process mixture models
- Variable selection in clustering via Dirichlet process mixture models
- Mixtures of products of Dirichlet processes for variable selection in survival analysis
- Fast approximation of variational Bayes Dirichlet process mixture using the maximization-maximization algorithm
- Bayesian Hierarchical Varying-Sparsity Regression Models with Application to Cancer Proteogenomics
- Bayesian approaches to variable selection in mixture models with application to disease clustering
- Variational inference for Dirichlet process mixtures
- Variable selection for sparse Dirichlet-multinomial regression with an application to microbiome data analysis
- Ultra-Fast Approximate Inference Using Variational Functional Mixed Models
- Bayesian curve fitting and clustering with Dirichlet process mixture models for microarray data
Cites work
- A Bayesian analysis of some nonparametric problems
- A framework for feature selection in clustering
- Bayesian Density Estimation and Inference Using Mixtures
- Bayesian model averaging: A tutorial. (with comments and a rejoinder).
- Bayesian Variable Selection in Clustering High-Dimensional Data
- Comparison of Discrimination Methods for the Classification of Tumors Using Gene Expression Data
- Estimating Normal Means with a Dirichlet Process Prior
- Estimating the dimension of a model
- Ferguson distributions via Polya urn schemes
- Hierarchical Dirichlet Processes
- scientific article; zbMATH DE number 3942813 (Why is no real title available?)
- scientific article; zbMATH DE number 720689 (Why is no real title available?)
- scientific article; zbMATH DE number 3046453 (Why is no real title available?)
- Improved criteria for clustering based on the posterior similarity matrix
- Mixtures of Dirichlet processes with applications to Bayesian nonparametric problems
- Model Selection and Accounting for Model Uncertainty in Graphical Models Using Occam's Window
- Model-Based Clustering, Discriminant Analysis, and Density Estimation
- On a class of Bayesian nonparametric estimates: I. Density estimates
- Prior distributions on spaces of probability measures
- Simple approximate MAP inference for Dirichlet processes mixtures
- Variable Selection for Clustering with Gaussian Mixture Models
- Variable Selection for Model-Based Clustering
- Variable selection for model-based clustering using the integrated complete-data likelihood
- Variable selection in clustering via Dirichlet process mixture models
- Variable selection methods for model-based clustering
- Variational inference for Dirichlet process mixtures
Cited in
(4)- ParticleMDI: particle Monte Carlo methods for the cluster analysis of multiple datasets with applications to cancer subtype identification
- Collocation based training of neural ordinary differential equations
- A comparative review of variable selection techniques for covariate dependent Dirichlet process mixture models
- Annealed variational mixtures for disease subtyping and biomarker discovery
This page was built for publication: Fast approximate inference for variable selection in Dirichlet process mixtures, with an application to pan-cancer proteomics
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2195280)