Abstract: This is an up-to-date introduction to and overview of the Minimum Description Length (MDL) Principle, a theory of inductive inference that can be applied to general problems in statistics, machine learning and pattern recognition. While MDL was originally based on data compression ideas, this introduction can be read without any knowledge thereof. It takes into account all major developments since 2007, the last time an extensive overview was written. These include new methods for model selection and averaging and hypothesis testing, as well as the first completely general definition of {em MDL estimators}. Incorporating these developments, MDL can be seen as a powerful extension of both penalized likelihood and Bayesian approaches, in which penalization functions and prior distributions are replaced by more general luckiness functions, average-case methodology is replaced by a more robust worst-case approach, and in which methods classically viewed as highly distinct, such as AIC vs BIC and cross-validation vs Bayes can, to a large extent, be viewed from a unified perspective.
Recommendations
Cites work
- A linear-time algorithm for computing the multinomial stochastic complexity
- A widely applicable Bayesian information criterion
- Achievability of asymptotic minimax regret by horizon-dependent and horizon-independent strategies
- Almost the best of three worlds: risk, consistency and optional stopping for the switch criterion in nested model selection
- An Information Measure for Classification
- Asymptotic equivalence of Bayes cross validation and widely applicable information criterion in singular learning theory
- Bayes Factors
- Bayes factors and marginal distributions in invariant situations
- Can the strengths of AIC and BIC be shared? A conflict between model indentification and regression estimation
- Catching up Faster by Switching Sooner: A Predictive Approach to Adaptive Estimation with an Application to the AIC–BIC Dilemma
- Counting probability distributions: Differential geometry and model selection
- Efficient Computation of Normalized Maximum Likelihood Codes for Gaussian Mixture Models With Its Applications to Clustering
- Finite-Sample Risk Bounds for Maximum Likelihood Estimation With Arbitrary Penalties
- Fisher information and stochastic complexity
- Flat Minima
- From -entropy to KL-entropy: analysis of minimum information complexity density estima\-tion
- Harold Jeffreys's default Bayes factor hypothesis tests: explanation, extension, and application in psychology
- High-dimensional penalty selection via minimum description length principle
- scientific article; zbMATH DE number 4095371 (Why is no real title available?)
- scientific article; zbMATH DE number 45100 (Why is no real title available?)
- scientific article; zbMATH DE number 107482 (Why is no real title available?)
- scientific article; zbMATH DE number 2080456 (Why is no real title available?)
- Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it
- Information and complexity in statistical modeling.
- Information-theoretic upper and lower bounds for statistical estimation
- Jeffreys versus Shtarkov distributions associated with some natural exponential families
- Krimp: mining itemsets that compress
- Learning Bayesian networks: The combination of knowledge and statistical data
- Learning Theory
- Minimax optimal Bayes mixtures for memoryless sources over large alphabets
- Minimax Pointwise Redundancy for Memoryless Models Over Large Alphabets
- Minimum complexity density estimation
- Minimum description length induction, Bayesianism, and Kolmogorov complexity
- Model selection by sequentially normalized least squares
- Modeling by shortest data description
- On model selection, Bayesian networks, and the Fisher information integral
- PAC-Bayesian stochastic model selection
- PAC-MDL bounds.
- Prediction, Learning, and Games
- Present Position and Potential Developments: Some Personal Views: Statistical Theory: The Prequential Approach
- Probabilistic graphical models.
- Relations Between the Conditional Normalized Maximum Likelihood Distributions and the Latent Information Priors
- Safe probability
- Subset selection in linear regression using sequentially normalized least squares: asymptotic theory
- Summarizing and understanding large graphs
- The elements of statistical learning. Data mining, inference, and prediction
- The geometry of proper scoring rules
- The minimum description length principle in coding and modeling
- The safe Bayesian. Learning the learning rate via the mixability gap
- Understanding machine learning. From theory to algorithms
- Unified Conditional Frequentist and Bayesian Testing of Composite Hypotheses
- Universal coding, information, prediction, and estimation
Cited in
(44)- The minimum description length principle for pattern mining: a survey
- Robust subgroup discovery. Discovering subgroup lists using MDL
- Minimum Description Length Principle for Fat-Tailed Distributions
- Information and complexity in statistical modeling.
- scientific article; zbMATH DE number 4166891 (Why is no real title available?)
- scientific article; zbMATH DE number 740677 (Why is no real title available?)
- The whole and the parts: the minimum description length principle and the a-contrario framework
- Minimum description length principle for linear mixed effects models
- scientific article; zbMATH DE number 6751371 (Why is no real title available?)
- scientific article; zbMATH DE number 2188025 (Why is no real title available?)
- scientific article; zbMATH DE number 4187170 (Why is no real title available?)
- Marginal Likelihood Computation for Model Selection and Hypothesis Testing: An Extensive Review
- Explanatory and creative alternatives to the MDL principle
- Unsupervised discretization by two-dimensional MDL-based histogram
- The no-free-lunch theorems of supervised learning
- On the parametric description of log-growth rates of Romanian city sizes
- Likelihood level adapted estimation of marginal likelihood for Bayesian model selection
- Authors' reply to the discussion of `safe testing'
- Zihao Wen and David L. Dowe's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Maozai Tian, Keming Yu and Jiangfeng Wang's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Andrej Srakar's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Judith ter Schure's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Christian P. Robert and Joshua Bon's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Stefano Rizzelli's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Luigi Pace and Alessandra Salvan's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Joris Mulder's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Alexander Ly's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Sander Greenland's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Neil Dey, Ryan Martin, and Jonathan P. Williams' contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Christine P. Chai's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Marco Cattaneo's contribution to the discussion of ``Safe testing by Grünwald, de Heide, and Koolen
- Joshua bon and christian P. Robert's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Vladimir Vovk's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Wenkai Xu's contribution to the discussion of `safe testing' by Grünwald, de Heide and Koolen
- Christian Hennig's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Glenn Shafer's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Thorsten Dickhaus's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Martin larsson, aaditya ramdas, and johannes Ruf's contribution to the discussion of `safe testing' by Grünwald, de heide, and koolen
- David R. Bickel's contribution to the discussion of `safe testing' by Grünwald, de Heide, and Koolen
- Seconder of the vote of thanks to Grünwald, de Heide, and Koolen and contribution to the discussion of `safe testing'
- Proposer of the vote of thanks to Grünwald, de Heide, and Koolen and contribution to the discussion of `safe testing'
- Safe testing
- An extended minimum description length (MDL) penalty function for binomial proportions
- A geometric modeling of Occam's razor in deep learning
This page was built for publication: Minimum description length revisited
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q4997077)