Local multiple alignment via subgraph enumeration
blockingchainingflatteningmaximal cliquemolecular biologymultiple alignmentssequence comparisonsubgraph enumeration
Enumeration in graph theory (05C30) Extremal problems in graph theory (05C35) Applications of graph theory (05C90) Graph theory (including graph drawing) in computer science (68R10) Software, source code, etc. for problems pertaining to biology (92-04) Computational methods for problems pertaining to biology (92-08) Biochemistry, molecular biology (92C40)
The algorithms described in this paper were developed to solve the following sort of computational problem from molecular biology. Suppose a newly sequenced protein has been used as the query for a database search, and that statistically significant similarity has been found to, say, six sequences from the database. The problem now is to identify the motif or motifs that are common to most or all of the seven sequences. In particular, the following difficulties may arise. (1) A motif may not appear in all of the sequences. (2) Two distinct motifs may appear in different subsets of the sequences. (3) A motif may have multiple appearance in a sequence. (4) Two motifs may appear in different multiplicites and relative orders in two sequences. Moreover, as in most sequence analysis problems in biology, when we say that a motif appears in some sequences, we mean that there are substrings that approximately match in a sense that needs to be made precise. Finally, the alignment-scoring scheme most appropriate for providing the desired precision to the notion of ``approximately match may depend on both the set of sequences and the particular motif.NEWLINENEWLINENEWLINEWe discuss three problems, which we call blocking, chaining and flattening, that arise when computing a multiple-sequence alignment from given pairwise alignments. Blocking is the construction of gap-free multiple alignments, each called a ``block, from the pairwise alignments; it is formalized here as the enumeration of maximal cliques in a certain graph. Chaining is the identification of a collection of blocks that can appear together in a multiple alignment, which we formalize as determining a maximal connected subgraph (of a different graph) that satisfies certain consistency conditions. Flattening is the introduction of gaps within a chain of blocks to create a multiple alignment, which involves solving a problem of dynamic bipartite matching. For each problem, practical algorithms are presented and shown to be effective for analyzing sequences containing internal repeats.
- Multiple Alignment, Communication Cost, and Graph Matching
- Partially local multi-way alignments
- On graph-based data structures to multiple genome alignment
- On the complexity of sequence to graph alignment
- An Eulerian path approach to local multiple alignment for DNA sequences
- Computational and Information Science
- On the number of many-to-many alignments of multiple sequences
- scientific article; zbMATH DE number 2086217
- A New Algorithm for Generating All the Maximal Independent Sets
- A time-efficient, linar-space local similarity algorithm
- Algorithm 457: finding all cliques of an undirected graph
- An $n^{5/2} $ Algorithm for Maximum Matchings in Bipartite Graphs
- Efficient Algorithms for Listing Combinatorial Structures
- Efficient methods for multiple sequence alignment with guaranteed error bounds
- Generating All Maximal Independent Sets: NP-Hardness and Polynomial-Time Algorithms
- scientific article; zbMATH DE number 5542185 (Why is no real title available?)
- scientific article; zbMATH DE number 3639144 (Why is no real title available?)
- scientific article; zbMATH DE number 910858 (Why is no real title available?)
- Linear-space algorithms that build local alignments from fragments
- Multiple sequence comparison and consistency on multipartite graphs
- On generating all maximal independent sets
- Sparse dynamic programming II
- Multiple sequence comparison and consistency on multipartite graphs
- Partially local multi-way alignments
- Neighborhood functions and hill-climbing strategies dedicated to the generalized ungapped local multiple alignment
- Chaining algorithms for multiple genome comparison
- scientific article; zbMATH DE number 1615272 (Why is no real title available?)
- Multiple Alignment, Communication Cost, and Graph Matching
- scientific article; zbMATH DE number 1746434 (Why is no real title available?)
- scientific article; zbMATH DE number 2087049 (Why is no real title available?)
- Nonoverlapping local alignments (weighted independent sets of axis parallel rectangles)
- Multiple alignment of biological sequences with gap flexibility
- CompBioNet 2004: algorithms \& computational methods for biochemical and evolutionary networks. Proceedings of the conference, Recife, Brazil, December 15--18, 2004.
- Comparative Genomics
- A heuristic algorithm for multiple sequence alignment based on blocks
This page was built for publication: Local multiple alignment via subgraph enumeration
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5961633)