Bayesian feature selection with strongly regularizing priors maps to the Ising model
From MaRDI portal
Publication:5380347
Abstract: Identifying small subsets of features that are relevant for prediction and/or classification tasks is a central problem in machine learning and statistics. The feature selection task is especially important, and computationally difficult, for modern datasets where the number of features can be comparable to, or even exceed, the number of samples. Here, we show that feature selection with Bayesian inference takes a universal form and reduces to calculating the magnetizations of an Ising model, under some mild conditions. Our results exploit the observation that the evidence takes a universal form for strongly-regularizing priors --- priors that have a large effect on the posterior probability even in the infinite data limit. We derive explicit expressions for feature selection for generalized linear models, a large class of statistical techniques that include linear and logistic regression. We illustrate the power of our approach by analyzing feature selection in a logistic regression-based classifier trained to distinguish between the letters B and D in the notMNIST dataset.
Recommendations
- Structural, Syntactic, and Statistical Pattern Recognition
- High-dimensional Ising model selection with Bayesian information criteria
- High-dimensional Ising model selection using \(\ell _{1}\)-regularized logistic regression
- scientific article; zbMATH DE number 5957308
- Bayesian feature selection for classification with possibly large number of classes
Cites work
- scientific article; zbMATH DE number 2061729 (Why is no real title available?)
- scientific article; zbMATH DE number 1865509 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- A significance test for the lasso
- Parametric inference in the large data limit using maximally informative models
- Pattern recognition and machine learning.
- Probabilistic data modelling with adaptive TAP mean-field theory
- Regularization and Variable Selection Via the Elastic Net
- Statistical Inference, Occam's Razor, and Statistical Mechanics on the Space of Probability Distributions
Cited in
(2)
This page was built for publication: Bayesian feature selection with strongly regularizing priors maps to the Ising model
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5380347)