A Bayesian framework for molecular strain identification from mixed diagnostic samples
From MaRDI portal
Abstract: We provide a mathematical formulation and develop a computational framework for identifying multiple strains of microorganisms from mixed samples of DNA. Our method is applicable in public health domains where efficient identification of pathogens is paramount, e.g., for the monitoring of disease outbreaks. We formulate strain identification as an inverse problem that aims at simultaneously estimating a binary matrix (encoding presence or absence of mutations in each strain) and a real-valued vector (representing the mixture of strains) such that their product is approximately equal to the measured data vector. The problem at hand has a similar structure to blind deconvolution, except for the presence of binary constraints, which we enforce in our approach. Following a Bayesian approach, we derive a posterior density. We present two computational methods for solving the non-convex maximum a posteriori estimation problem. The first one is a local optimization method that is made efficient and scalable by decoupling the problem into smaller independent subproblems, whereas the second one yields a global minimizer by converting the problem into a convex mixed-integer quadratic programming problem. The decoupling approach also provides an efficient way to integrate over the posterior. This provides useful information about the ambiguity of the underdetermined problem and, thus, the uncertainty associated with numerical solutions. We evaluate the potential and limitations of our framework in silico using synthetic and experimental data with available ground truths.
Recommendations
- Magnitude and sources of bias in the detection of mixed strain \textit{M. tuberculosis} infection
- Non-uniqueness in probabilistic numerical identification of bacteria
- Computational aspects of DNA mixture analysis
- On inferring presence of an individual in a mixture: a Bayesian approach
- A deconvolution path for mixtures
Cites work
- A clustering approach for the blind separation of multiple finite alphabet sequences from a single linear mixture
- Computability of global solutions to factorable nonconvex programs: Part I — Convex underestimating problems
- Convergence of the alternating minimization algorithm for blind deconvolution
- Discrete inverse problems. Insight and algorithms.
- scientific article; zbMATH DE number 5060482 (Why is no real title available?)
- Identifiability for Blind Source Separation of Multiple Finite Alphabet Linear Mixtures
- Introduction to Bayesian Scientific Computing
- Julia: a fresh approach to numerical computing
- JuMP: a modeling language for mathematical optimization
- Mixed-integer nonlinear optimization
- Nonstationary inverse problems and state estimation
- Probability essentials.
- Projected Gradient Methods for Nonnegative Matrix Factorization
- Regularization by discretization in Banach spaces
- Relaxations and discretizations for the pooling problem
- Solving mixed integer bilinear problems using MILP formulations
Cited in
(2)
This page was built for publication: A Bayesian framework for molecular strain identification from mixed diagnostic samples
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4582733)