Modeling overdispersion heterogeneity in differential expression analysis using mixtures
From MaRDI portal
Abstract: Next-generation sequencing technologies now constitute a method of choice to measure gene expression. Data to analyze are read counts, commonly modeled using Negative Binomial distributions. A relevant issue associated with this probabilistic framework is the reliable estimation of the overdispersion parameter, reinforced by the limited number of replicates generally observable for each gene. Many strategies have been proposed to estimate this parameter, but when differential analysis is the purpose, they often result in procedures based on plug-in estimates, and we show here that this discrepancy between the estimation framework and the testing framework can lead to uncontrolled type-I errors. Instead we propose a mixture model that allows each gene to share information with other genes that exhibit similar variability. Three consistent statistical tests are developed for differential expression analysis. We show that the proposed method improves the sensitivity of detecting differentially expressed genes with respect to the common procedures, since it is the best one in reaching the nominal value for the first-type error, while keeping elevate power. The method is finally illustrated on prostate cancer RNA-seq data.
Recommendations
- The NBP negative binomial model for assessing differential gene expression from RNA-Seq
- BNP-seq: Bayesian nonparametric differential expression analysis of sequencing count data
- Detecting differential gene expression with a semiparametric hierarchical mixture method
- Detecting differential expression in RNA-sequence data using quasi-likelihood with shrunken dispersion estimates
- Mixture Model on the Variance for the Differential Analysis of Gene Expression Data
Cites work
- A two-stage Poisson model for testing RNA-Seq data
- Asymptotic Statistics
- DESeq2
- Detecting differential expression in RNA-sequence data using quasi-likelihood with shrunken dispersion estimates
- Finite mixture models
- scientific article; zbMATH DE number 3567782 (Why is no real title available?)
- scientific article; zbMATH DE number 720689 (Why is no real title available?)
- Mixture Model on the Variance for the Differential Analysis of Gene Expression Data
- Model-Based Clustering, Discriminant Analysis, and Density Estimation
- Small-sample estimation of negative binomial dispersion, with applications to SAGE data
- The NBP negative binomial model for assessing differential gene expression from RNA-Seq
Cited in
(5)- Single-gene negative binomial regression models for RNA-seq data with higher-order asymptotic inference
- A mixture modeling framework for differential analysis of high-throughput data
- Shrinkage of dispersion parameters in the binomial family, with application to differential exon skipping
- BNP-seq: Bayesian nonparametric differential expression analysis of sequencing count data
- Bayesian negative binomial mixture regression models for the analysis of sequence count and methylation data
Describes a project that uses
Uses Software
This page was built for publication: Modeling overdispersion heterogeneity in differential expression analysis using mixtures
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2827191)