Fast and fully-automated histograms for large-scale data sets
From MaRDI portal
Abstract: G-Enum histograms are a new fast and fully automated method for irregular histogram construction. By framing histogram construction as a density estimation problem and its automation as a model selection task, these histograms leverage the Minimum Description Length principle (MDL) to derive two different model selection criteria. Several proven theoretical results about these criteria give insights about their asymptotic behavior and are used to speed up their optimisation. These insights, combined to a greedy search heuristic, are used to construct histograms in linearithmic time rather than the polynomial time incurred by previous works. The capabilities of the proposed MDL density estimation method are illustrated with reference to other fully automated methods in the literature, both on synthetic and large real-world data sets.
Recommendations
Cites work
- A comparison of automatic histogram constructions
- A linear-time algorithm for computing the multinomial stochastic complexity
- A universal prior for integers and estimation by minimum description length
- Akaike's information criterion and Kullback-Leibler loss for histogram density estimation
- Akaike's information criterion and the histogram
- An optimal variable cell histogram
- Combining regular and irregular histograms by penalized likelihood
- Densities, spectral densities and modality.
- Density estimation by stochastic complexity
- Estimating the dimension of a model
- How many bins should be put in a regular histogram
- scientific article; zbMATH DE number 4095371 (Why is no real title available?)
- scientific article; zbMATH DE number 3789676 (Why is no real title available?)
- Identifying excessively rounded or truncated data
- Modeling by shortest data description
- MODL: a Bayes optimal discretization method for continuous attributes
- Nonparametric density estimation by exact leave-\(p\)-out cross-validation
- On asymptotics of certain recurrences arising in universal coding
- On optimal and data-based histograms
- On stochastic complexity and nonparametric density estimation
- On the approximation of curves by line segments using dynamic programming
- On the histogram as a density estimator:L 2 theory
- Optimal cross-validation in density estimation with the \(L^{2}\)-loss
- Stochastic complexity and modeling
- Strong optimality of the normalized ML models as universal codes and information in data
- The Efficiency of Histogram-like Techniques for Database Query Optimization
- The Essential Histogram
This page was built for publication: Fast and fully-automated histograms for large-scale data sets
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6167056)