An algorithm for the principal component analysis of large data sets
From MaRDI portal
Abstract: Recently popularized randomized methods for principal component analysis (PCA) efficiently and reliably produce nearly optimal accuracy --- even on parallel processors --- unlike the classical (deterministic) alternatives. We adapt one of these randomized methods for use with data sets that are too large to be stored in random-access memory (RAM). (The traditional terminology is that our procedure works efficiently "out-of-core.") We illustrate the performance of the algorithm via several numerical examples. For example, we report on the PCA of a data set stored on disk that is so large that less than a hundredth of it can fit in our computer's RAM.
Recommendations
Cited in
(30)- Randomized algorithms for distributed computation of principal component analysis and singular value decomposition
- Randomized LU decomposition using sparse projections
- Randomized block Krylov methods for approximating extreme eigenvalues
- Single-pass randomized QLP decomposition for low-rank approximation
- Functional principal subspace sampling for large scale functional data analysis
- Model order reduction for nonlinear Schrödinger equation
- Algorithm 971
- A randomized algorithm for principal component analysis
- A stochastic variance reduction method for PCA by an exact penalty approach
- Principal Component Analysis of Large Dispersion Matrices
- scientific article; zbMATH DE number 839297 (Why is no real title available?)
- Gradient algorithms for principal component analysis
- Summation pollution of principal component analysis and an improved algorithm for location sensitive data
- Approximating matrix eigenvalues by subspace iteration with repeated random sparsification
- Compactification of the rigid motions group in image processing
- Fast deflation sparse principal component analysis via subspace projections
- Pass-efficient randomized algorithms for low-rank matrix approximation using any number of views
- Streaming low-rank matrix approximation with an application to scientific simulation
- Fast and Accurate Proper Orthogonal Decomposition using Efficient Sampling and Iterative Techniques for Singular Value Decomposition
- Randomized numerical linear algebra: Foundations and algorithms
- Robust Recovery of Low-Rank Matrices and Low-Tubal-Rank Tensors from Noisy Sketches
- Deep learning methods for partial differential equations and related parameter identification problems
- Krylov-Aware Stochastic Trace Estimation
- Principal Component Analysis and Randomness Test for Big Data Analysis
- Fast and accurate randomized algorithms for linear systems and eigenvalue problems
- Solution of large linear discrete ill-posed problems by randomized block Krylov methods
- A generalized Nyström method with subspace iteration for low-rank approximations of large-scale nonsymmetric matrices
- Estimation of local geometric structure on manifolds from noisy data
- An improved algorithm for the RBFNN based on SVD
- Randomized block-Krylov subspace methods for low-rank approximation of matrix functions
This page was built for publication: An algorithm for the principal component analysis of large data sets
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3116449)