Bayesian approximate kernel regression with variable selection
From MaRDI portal
Abstract: Nonlinear kernel regression models are often used in statistics and machine learning because they are more accurate than linear models. Variable selection for kernel regression models is a challenge partly because, unlike the linear regression setting, there is no clear concept of an effect size for regression coefficients. In this paper, we propose a novel framework that provides an effect size analog of each explanatory variable for Bayesian kernel regression models when the kernel is shift-invariant --- for example, the Gaussian kernel. We use function analytic properties of shift-invariant reproducing kernel Hilbert spaces (RKHS) to define a linear vector space that: (i) captures nonlinear structure, and (ii) can be projected onto the original explanatory variables. The projection onto the original explanatory variables serves as an analog of effect sizes. The specific function analytic property we use is that shift-invariant kernel functions can be approximated via random Fourier bases. Based on the random Fourier expansion we propose a computationally efficient class of Bayesian approximate kernel regression (BAKR) models for both nonlinear regression and binary classification for which one can compute an analog of effect sizes. We illustrate the utility of BAKR by examining two important problems in statistical genetics: genomic selection (i.e. phenotypic prediction) and association mapping (i.e. inference of significant variants or loci). State-of-the-art methods for genomic selection and association mapping are based on kernel regression and linear models, respectively. BAKR is the first method that is competitive in both settings.
Recommendations
- Variable selection for nonparametric learning with power series kernels
- Model selection in kernel ridge regression
- Selection of the bandwidth parameter in a Bayesian kernel regression model for genomic-enabled prediction
- Scalable variational inference for Bayesian variable selection in regression, and its accuracy in genetic association studies
- scientific article; zbMATH DE number 1036034
Cites work
- 10.1162/153244303322753706
- A Brief Survey of Bandwidth Selection for Density Estimation
- Bayesian binary kernel probit model for microarray based cancer classification and gene selection
- Bayesian Classification of Tumours by Using Gene Expression Data
- Bayesian generalized kernel mixed models
- Bayesian nonlinear regression for large \(p\) small \(n\) problems
- Characterizing the function space for Bayesian kernel models
- Choosing multiple parameters for support vector machines
- Gaussian processes for machine learning.
- Gene expression-based glioma classification using hierarchical Bayesian vector machines
- scientific article; zbMATH DE number 1804115 (Why is no real title available?)
- scientific article; zbMATH DE number 45848 (Why is no real title available?)
- scientific article; zbMATH DE number 6860845 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- LDAK
- Least angle regression. (With discussion)
- Nonparametric sparsity and regularization
- On the equivalence between kernel quadrature rules and random feature expansions
- Partial Factor Modeling: Predictor-Dependent Shrinkage for Linear Regression
- Scalable Bayesian nonparametric regression via a Plackett-Luce model for conditional ranks
- The Bayesian Lasso
- The elements of statistical learning. Data mining, inference, and prediction
Cited in
(15)- Convex and non-convex regularization methods for spatial point processes intensity estimation
- Nowcasting in a pandemic using non-parametric mixed frequency VARs
- Sparse Bayesian variable selection in kernel probit model for analyzing high-dimensional data
- A statistical pipeline for identifying physical features that differentiate classes of 3D shapes
- Variable prioritization in nonlinear black box methods: a genetic association case study
- scientific article; zbMATH DE number 5968941 (Why is no real title available?)
- Model Interpretation Through Lower-Dimensional Posterior Summarization
- The Holdout Randomization Test for Feature Selection in Black Box Models
- Predicting clinical outcomes in glioblastoma: an application of topological and functional data analysis
- TAIL FORECASTING WITH MULTIVARIATE BAYESIAN ADDITIVE REGRESSION TREES
- Non-linear dimension reduction in factor-augmented vector autoregressions
- A simple approach for local and global variable importance in nonlinear regression models
- A novel power-based approach to Gaussian kernel selection in the kernel-based association test
- Nonparametric mixed frequency monitoring macro-at-risk
- Selection of the bandwidth parameter in a Bayesian kernel regression model for genomic-enabled prediction
This page was built for publication: Bayesian approximate kernel regression with variable selection
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3121562)