Deep learning: a statistical viewpoint
From MaRDI portal
Recommendations
- scientific article; zbMATH DE number 7387620
- A contribution to the statistical theory of deep learning
- Understanding Deep Learning with Statistical Relevance
- Deep learning: a Bayesian perspective
- Machine learning: deepest learning as statistical data assimilation problems
- Deep neural networks for estimation and inference
- A statistician teaches deep learning
- scientific article; zbMATH DE number 6975079
- scientific article; zbMATH DE number 1413009
- Deep learning: methods and applications
Cites work
- 10.1162/153244302760200704
- 10.1162/153244303321897690
- A better algorithm for random \(k\)-SAT
- A decision-theoretic generalization of on-line learning and an application to boosting
- A general lower bound on the number of examples needed for learning
- A jamming transition from under- to over-parametrization affects generalization in deep learning
- A mean field view of the landscape of two-layer neural networks
- A note on margin-based loss functions in classification
- A random matrix approach to neural networks
- A risk comparison of ordinary least squares vs ridge regression
- An inequality for uniform deviations of sample averages from their means
- Analysis of Two Simple Heuristics on a Random Instance ofk-sat
- Anisotropic local laws for random matrices
- Arcing classifiers. (With discussion)
- Benign overfitting in linear regression
- Boosting the margin: a new explanation for the effectiveness of voting methods
- Boosting with early stopping: convergence and consistency
- Bounding the smallest singular value of a random matrix without concentration
- Bounds on rates of variable-basis and neural-network approximation
- Breaking the curse of dimensionality with convex neural networks
- Combinatorics of random processes and sections of convex bodies
- Comparison of worst case errors in linear and neural network approximation
- Concentration inequalities and moment bounds for sample covariance operators
- Concentration inequalities. A nonasymptotic theory of independence
- Convexity, Classification, and Risk Bounds
- Decision theoretic generalizations of the PAC model for neural net and other learning applications
- Deep learning
- Distribution-free inequalities for the deleted and holdout error estimates
- Does learning require memorization? a short tale about a long tail
- Efficient agnostic learning of neural networks with bounded fan-in
- Empirical entropy, minimax regret and minimax risk
- Estimating a regression function
- Extending the scope of the small-ball method
- FAST RATES FOR ESTIMATION ERROR AND ORACLE INEQUALITIES FOR MODEL SELECTION
- Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration
- Gradient descent optimizes over-parameterized deep ReLU networks
- Gradient flows in metric spaces and in the space of probability measures
- Greedy function approximation: A gradient boosting machine.
- Hardness results for neural network approximation problems
- High-dimensional probability. An introduction with applications in data science
- scientific article; zbMATH DE number 49190 (Why is no real title available?)
- scientific article; zbMATH DE number 51427 (Why is no real title available?)
- scientific article; zbMATH DE number 3626409 (Why is no real title available?)
- scientific article; zbMATH DE number 3639144 (Why is no real title available?)
- scientific article; zbMATH DE number 1182753 (Why is no real title available?)
- scientific article; zbMATH DE number 2038320 (Why is no real title available?)
- scientific article; zbMATH DE number 1552503 (Why is no real title available?)
- scientific article; zbMATH DE number 3446442 (Why is no real title available?)
- scientific article; zbMATH DE number 819196 (Why is no real title available?)
- scientific article; zbMATH DE number 845714 (Why is no real title available?)
- scientific article; zbMATH DE number 7626719 (Why is no real title available?)
- scientific article; zbMATH DE number 6781369 (Why is no real title available?)
- scientific article; zbMATH DE number 7064043 (Why is no real title available?)
- scientific article; zbMATH DE number 3222478 (Why is no real title available?)
- Improving the sample complexity using global data
- Introduction to nonparametric estimation
- Just interpolate: kernel ``ridgeless regression can generalize
- Kernels as features: on kernels, margins, and low-dimensional mappings
- Learnability and the Vapnik-Chervonenkis dimension
- Linearized two-layers neural networks in high dimension
- Local Rademacher complexities
- Local Rademacher complexities and oracle inequalities in risk minimization. (2004 IMS Medallion Lecture). (With discussions and rejoinder)
- Mean field analysis of neural networks: a law of large numbers
- Model selection and error estimation
- Nearest neighbor pattern classification
- Neural Network Learning
- Nonlinear random matrix theory for deep learning
- On the Bayes-risk consistency of regularized boosting methods.
- On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities
- Optimal rates for the regularized least-squares algorithm
- Optimal transport for applied mathematicians. Calculus of variations, PDEs, and modeling
- Polynomial bounds for VC dimension of sigmoidal and general Pfaffian neural networks
- Products of many large random matrices and gradients in deep neural networks
- Rademacher penalties and structural risk minimization
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- Sharper bounds for Gaussian and empirical processes
- Size-independent sample complexity of neural networks
- Statistical behavior and consistency of classification methods based on convex risk minimization.
- Support-vector networks
- Surprises in high-dimensional ridgeless least squares interpolation
- The concentration of measure phenomenon
- The densest hemisphere problem
- The elements of statistical learning. Data mining, inference, and prediction
- The Generalization Error of Random Features Regression: Precise Asymptotics and the Double Descent Curve
- The Hilbert kernel regression estimate.
- The implicit bias of gradient descent on separable data
- The interpolation phase transition in neural networks: memorization and generalization under lazy training
- The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
- The spectral norm of random inner-product kernel matrices
- The spectrum of kernel random matrices
- The spectrum of random inner-product kernel matrices
- The spectrum of random kernel matrices: universality results for rough and varying kernels
- Theoretical foundations of the potential function method in pattern recognition learning
- When do neural networks outperform kernel methods?*
Cited in
(62)- The interpolation phase transition in neural networks: memorization and generalization under lazy training
- From inexact optimization to learning via gradient concentration
- Surprises in high-dimensional ridgeless least squares interpolation
- On the minimax optimality and superiority of deep neural network learning over sparse parameter spaces
- Free dynamics of feature learning processes
- Should we estimate a product of density functions by a product of estimators?
- scientific article; zbMATH DE number 6975079 (Why is no real title available?)
- Learning Summary Statistic for Approximate Bayesian Computation via Deep Neural Network
- The implicit bias of gradient descent on separable data
- Two models of double descent for weak features
- Generalization in Overparameterized Models
- Binary classification of Gaussian mixtures: abundance of support vectors, benign overfitting, and regularization
- Theoretical issues in deep networks
- Benign overfitting in linear regression
- Optimal regularizations for data generation with probabilistic graphical models
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- Scaling description of generalization with number of parameters in deep learning
- The Modern Mathematics of Deep Learning
- Generalization in Deep Learning
- Understanding Deep Learning with Statistical Relevance
- Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
- A note on the prediction error of principal component regression in high dimensions
- Adversarial examples in random neural networks with general activations
- Statistical guarantees for regularized neural networks
- Mini-workshop: Mathematical foundations of robust and generalizable learning. Abstracts from the mini-workshop held October 2--8, 2022
- High-Dimensional Analysis of Double Descent for Linear Regression with Random Projections
- Measuring Complexity of Learning Schemes Using Hessian-Schatten Total Variation
- Tractability from overparametrization: the example of the negative perceptron
- More is Less: Inducing Sparsity via Overparameterization
- A moment-matching approach to testable learning and a new characterization of Rademacher complexity
- Mini-workshop: Interpolation and over-parameterization in statistics and machine learning. Abstracts from the mini-workshop held September 17--22, 2023
- The curse of overparametrization in adversarial training: precise analysis of robust generalization for random features regression
- Out-of-distributional risk bounds for neural operators with applications to the Helmholtz equation
- Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks
- Convergence analysis for over-parameterized deep learning
- Learning sparse features can lead to overfitting in neural networks
- Mini-workshop: Nonlinear approximation of high-dimensional functions in scientific computing. Abstracts from the mini-workshop held October 15--20, 2023
- Double data piling: a high-dimensional solution for asymptotically perfect multi-category classification
- Deep networks for system identification: a survey
- Neural networks generalize on low complexity data
- Analysis of the expected L₂ error of an over-parametrized deep neural network estimate learned by gradient descent without regularization
- A geometrical analysis of kernel ridge regression and its applications
- DRM revisited: a complete error analysis
- Learning time-scales in two-layers neural networks
- Corrected generalized cross-validation for finite ensembles of penalized estimators
- On the rate of convergence of an over-parametrized deep neural network regression estimate with ReLU activation function learned by gradient descent
- Consistency and rate of convergence for deep ReLU neural networks
- Convergence analysis of PINNs with over-parameterization
- The positivity of the neural tangent kernel
- Quantitative CLTs in deep neural networks
- Convergence and recovery guarantees of unsupervised neural networks for inverse problems
- Random features and polynomial rules
- Dimension free ridge regression
- Tensor-on-tensor regression: Riemannian optimization, over-parameterization, statistical-computational gap and their interplay
- Precise asymptotics of bagging regularized M-estimators
- Six lectures on linearized neural networks
- Gaussian random projections of convex cones: approximate kinematic formulae and applications
- Universality of kernel random matrices and kernel regression in the quadratic regime
- Accelerating neural network-based regression and classification tasks through empirical kernels
- Sources of uncertainty in supervised machine learning -- a statisticians' view
- Local geometry of high-dimensional mixture models: effective spectral theory and dynamical transitions
- Sobolev norm inconsistency of kernel interpolation
This page was built for publication: Deep learning: a statistical viewpoint
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5887827)