Matching strings in encoded sequences

From MaRDI portal



Abstract: We investigate the longest common substring problem for encoded sequences and its asymptotic behaviour. The main result is a strong law of large numbers for a re-scaled version of this quantity, which presents an explicit relation with the R'enyi entropy of the source. We apply this result to the zero-inflated contamination model and the stochastic scrabble. In the case of dynamical systems, this problem is equivalent to the shortest distance between two observed orbits and its limiting relationship with the correlation dimension of the pushforward measure. An extension to the shortest distance between orbits for random dynamical systems is also provided.


Let \(\chi, \tilde\chi \) be alphabets. An encoder is a measurable function on strings \(f\colon \chi ^{\mathrm{N}} \to \tilde\chi^{\mathrm{N}}\). For any \(x,y \in \chi^{\mathrm{N}}\) the longest common substring between encoded strings is \[ M_n^f(x,y) = \max \{ k:f(x)_i^{i + k - 1} = f(y)_j^{j + k - 1},\ 0 \le i,j \le n - k\}. \] Here it is proved an almost sure convergence of \(M_n^f\). The result is applied to generalize some earlier results for stochastic scrabble, stochastic noise and for analysis of the behaviour of the shortest distance between observed orbits of a dynamical system.



Cites work









This page was built for publication: Matching strings in encoded sequences

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2174991)