Sublinear algorithms for approximating string compressibility
From MaRDI portal
Publication:2392931
DOI10.1007/S00453-012-9618-6zbMATH Open1270.68109arXiv0706.1084OpenAlexW2173739051MaRDI QIDQ2392931FDOQ2392931
Authors: Yanyan Li
Publication date: 5 August 2013
Published in: Algorithmica (Search for Journal in Brave)
Abstract: We raise the question of approximating the compressibility of a string with respect to a fixed compression scheme, in sublinear time. We study this question in detail for two popular lossless compression schemes: run-length encoding (RLE) and Lempel-Ziv (LZ), and present sublinear algorithms for approximating compressibility with respect to both schemes. We also give several lower bounds that show that our algorithms for both schemes cannot be improved significantly. Our investigation of LZ yields results whose interest goes beyond the initial questions we set out to study. In particular, we prove combinatorial structural lemmas that relate the compressibility of a string with respect to Lempel-Ziv to the number of distinct short substrings contained in it. In addition, we show that approximating the compressibility with respect to LZ is related to approximating the support size of a distribution.
Full work available at URL: https://arxiv.org/abs/0706.1084
Recommendations
Nonnumerical algorithms (68W05) Coding and information theory (compaction, compression, models of communication, encoding schemes, etc.) (aspects in computer science) (68P30)
Cites Work
- Title not available (Why is that?)
- Title not available (Why is that?)
- Title not available (Why is that?)
- The space complexity of approximating the frequency moments
- Title not available (Why is that?)
- Compression of individual sequences via variable-rate coding
- Sampling algorithms: lower bounds and applications
- Clustering by Compression
- The Similarity Metric
- A universal algorithm for sequential data compression
- Estimation of Entropy and Mutual Information
- Generalized substring compression
- Substring compression problems
- Discrete Cosine Transform
- On average sequence complexity
- Title not available (Why is that?)
- The context-tree weighting method: basic properties
- The Complexity of Approximating the Entropy
- Universal Entropy Estimation Via Block Sorting
- Strong lower bounds for approximating distribution support size and the distinct elements problem
- On the combinatorics of finite words
- Using literal and grammatical statistics for authorship attribution
- Proof of a conjecture on word complexity
- Theory and Applications of Models of Computation
- On the maximum number of distinct factors of a binary string
- Title not available (Why is that?)
- Sublinear Algorithms for Approximating String Compressibility
- Title not available (Why is that?)
- Estimating Entropy on<tex>$m$</tex>Bins Given Fewer Than<tex>$m$</tex>Samples
- Approximating entropy from sublinear samples
- On correlation polynomials and subword complexity
Cited In (18)
- Upper bounds on distinct maximal (sub-)repetitions in compressed strings
- Compressed Dynamic Tries with Applications to LZ-Compression in Sublinear Time and Space
- On stricter reachable repetitiveness measures
- At the roots of dictionary compression: string attractors
- Near-optimal search time in \(\delta \)-optimal space
- Sublinear Algorithms for Approximating String Compressibility
- Internal shortest absent word queries in constant time and linear space
- Substring complexities on run-length compressed strings
- Sensitivity of string compressors and repetitiveness measures
- Adaptive learning of compressible strings
- Compressibility measures for two-dimensional data
- Frequency-constrained substring complexity
- Sublinear time Lempel-Ziv (LZ77) factorization
- Iterated straight-line programs
- CONCUR 2003 - Concurrency Theory
- Title not available (Why is that?)
- NC algorithms for finding a maximal set of paths with application to compressing strings
- Near-optimal search time in \(\delta \)-optimal space, and vice versa
This page was built for publication: Sublinear algorithms for approximating string compressibility
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2392931)