Large Sample Properties of Partitioning-Based Series Estimators

From MaRDI portal



Abstract: We present large sample results for partitioning-based least squares nonparametric regression, a popular method for approximating conditional expectation functions in statistics, econometrics, and machine learning. First, we obtain a general characterization of their leading asymptotic bias. Second, we establish integrated mean squared error approximations for the point estimator and propose feasible tuning parameter selection. Third, we develop pointwise inference methods based on undersmoothing and robust bias correction. Fourth, employing different coupling approaches, we develop uniform distributional approximations for the undersmoothed and robust bias-corrected t-statistic processes and construct valid confidence bands. In the univariate case, our uniform distributional approximations require seemingly minimal rate restrictions and improve on approximation rates known in the literature. Finally, we apply our general results to three partitioning-based estimators: splines, wavelets, and piecewise polynomials. The supplemental appendix includes several other general and example-specific technical and methodological results. A companion R package is provided.


The authors study nonparametric regression problems for univariate responses \(y_1, \ldots, y_n\) and \(\mathbb{R}^d\)-valued, continuously distributed covariates \(\mathbf{x}_1, \ldots, \mathbf{x}_n\), where the latter are supported on the compact set \(\mathcal{X}\). In this, the object of interest is the (mean) regression function \(\mu(\cdot)\), such that \(\mu(\mathbf{x}) = \mathbb{E}[y | \mathbf{x}]\). The authors consider partitioning-based series least squares estimators (LSEs) for \(\mu(\cdot)\) and its derivatives, meaning that \(\mathcal{X}\) is partitioned into non-overlapping cells, on which basis functions are defined. Examples are spline bases, compactly supported wavelet bases, and piecewise polynomial bases. First, the (asymptotic) bias of the LSE is characterized by means of its leading term. Based on this, three bias correction methods are derived. Second, the performance of the LSE is analyzed in terms of the asymptotic behaviour of its integrated mean squared error. Third, pointwise (for fixed \(\mathbf{x}\)) and uniform (over \(\mathcal{X}\)) inference methods are elaborated upon in terms of central limit theorems and strong approximations, respectively, based on undersmoothing and robust bias correction. For practical purposes, the authors also propose ways of feasible tuning parameter selection, and they illustrate their theoretical findings by means of Monte Carlo simulations.



Cites work



Describes a project that uses

Uses Software






This page was built for publication: Large Sample Properties of Partitioning-Based Series Estimators

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q143957)