On compound Poisson approximation for sequence matching
Let \(A_1,\dots, A_m\) and \(B_1,\dots, B_n\) each be independent and identically distributed sequences of random variables from distributions \(\mu\) and \(\nu\), respectively, over a finite alphabet. The paper considers the distribution of the number \(W\) of pairs \(i\), \(j\) such that \(A_{i+l}= B_{j+l}\) for all but at most \(r\) values of \(l\), \(0\leq l\leq k-1\), for \(k\) suitably large; the motivation is to approximate the null hypothesis distributions of statistics measuring the degree of local similarity between two molecular sequences, where some mismatches are allowed. The Stein-Chen method is used to establish the accuracy of a compound Poisson approximation to the distribution of \(W\) with respect to the total variation distance; the distance between the two is shown to be asymptotically negligible in a wide variety of settings.
- Estimate of the Accuracy of the Compound Poisson Approximation for the Distribution of the Number of Matching Patterns
- A Poisson approximation for sequence comparisons with insertions and deletions
- Poisson, compound Poisson and process approximations for testing statistical significance in sequence comparisons
- A Phase Transition for the Distribution of Matching Blocks
- scientific article; zbMATH DE number 850337
- Poisson, compound Poisson and process approximations for testing statistical significance in sequence comparisons
- A Poisson approximation for sequence comparisons with insertions and deletions
- Compound Poisson process approximation.
- Compound Poisson approximation for regularly varying fields with application to sequence alignment
- Longest common substring for random subshifts of finite type
- Matching strings in encoded sequences
- Compound Poisson approximation for long increasing sequences
- Poisson approximation and dna sequence matching
- Sequence comparison with mixed convex and concave costs
- Estimate of the Accuracy of the Compound Poisson Approximation for the Distribution of the Number of Matching Patterns
- Distributional regimes for the number of k -word matches between two random sequences
This page was built for publication: On compound Poisson approximation for sequence matching
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2711618)