Approximate word matches between two random sequences
Let \({\mathcal L}= \{A,G,C,T\}\) be an alphabet with strand-symmetric probability measure \(\xi= \{\xi_A, \xi_G, \xi_C, \xi_T\}\) and perturbation parameter \(\eta\). That is, \(-1\leq\eta\leq 1\) is the unique number satisfying \(\xi_A= \xi_T= (1/4)(1+ \eta)\), \(\xi_C= \xi_G= (1/4)(1- \eta)\). Let \({\mathbf A}= A_1 A_2\cdots A_n\) and \({\mathbf B}= B_1 B_2\cdots B_n\) be two random squences of length \(n\) over \({\mathcal L}\) such that the letters are i.i.d. Let \({\mathbf x}\) and \({\mathbf y}\) be two words of length \(m< n\) and define the distance between \({\mathbf x}\) and \({\mathbf y}\) by \[ \delta({\mathbf x},{\mathbf y})=\text{number of characters mismatches between \({\mathbf x}\) and \({\mathbf y}\)}. \] We say that \({\mathbf x}\) is a \(k\)-neighbor of \({\mathbf y}\) if \(\delta({\mathbf x},{\mathbf y})\leq k< m\). Furthermore, define the statistic \(D^{(k)}_2\) to be the number of \(k\)-neighborhood \(m\)-word matches between the sequences \({\mathbf A}\) and \({\mathbf B}\), including overlaps. In this paper, the authors consider the mean of \(D^{(k)}_2\) and a lower and upper bound for the variance of \(D^{(k)}_2\) and prove among others the following theorem: If \(\eta\neq 0\), \(m= \alpha\log_{1/p_2}(n)+ C\) with \(0\leq\alpha<{1\over 2}\) and \(C\) a constant and \(0\leq k< m\) is fixed, then \[ {D^{(k)}_2- E(D^{(k)}_2)\over \sqrt{\text{Var}(D^{(k)}_2)}}\overset {d}\Rightarrow{\mathcal N}(0,1)\quad\text{as }{n}\infty. \] Here, \(p_2= \sum_{a\in{\mathcal L}}\xi^2_a\). The results are useful in the study of bioinformatics for expressed sequence tag database searches.
- Asymptotic Behavior of k-Word Matches Between two Uniformly Distributed Sequences
- Distributional regimes for the number of k -word matches between two random sequences
- An extreme value theory for sequence matching
- Counts of long aligned word matches among random letter sequences
- An accurate approximation to the distribution of the length of the longest matching word between two random DNA sequences
- Asymptotic Behavior of k-Word Matches Between two Uniformly Distributed Sequences
- Compound Poisson approximation: A user's guide
- Distributional regimes for the number of k -word matches between two random sequences
- scientific article; zbMATH DE number 50805 (Why is no real title available?)
- scientific article; zbMATH DE number 3438144 (Why is no real title available?)
- scientific article; zbMATH DE number 1912144 (Why is no real title available?)
- scientific article; zbMATH DE number 850226 (Why is no real title available?)
- Normal convergence by higher semi-invariants with applications to sums of dependent random variables and random graphs
- Poisson approximation for dependent trials
- Empirical distribution of \(k\)-word matches in biological sequences
- Limit distributions of extremal distances to the nearest neighbor
- New powerful statistics for alignment-free sequence comparison under a pattern transfer model
- Indifference pricing for CRRA utilities
- Extraction of high quality \(k\)-words for alignment-free sequence comparison
- Statistical considerations underpinning an alignment-free sequence comparison method
- Counts of long aligned word matches among random letter sequences
- Distributional regimes for the number of k -word matches between two random sequences
- Scoring unusual words with varying mismatch errors
This page was built for publication: Approximate word matches between two random sequences
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2476396)