Bi-cross-validation for factor analysis
From MaRDI portal
Abstract: Factor analysis is over a century old, but it is still problematic to choose the number of factors for a given data set. The scree test is popular but subjective. The best performing objective methods are recommended on the basis of simulations. We introduce a method based on bi-cross-validation, using randomly held-out submatrices of the data to choose the number of factors. We find it performs better than the leading methods of parallel analysis (PA) and Kaiser's rule. Our performance criterion is based on recovery of the underlying factor-loading (signal) matrix rather than identifying the true number of factors. Like previous comparisons, our work is simulation based. Recent advances in random matrix theory provide principled choices for the number of factors when the noise is homoscedastic, but not for the heteroscedastic case. The simulations we choose are designed using guidance from random matrix theory. In particular, we include factors too small to detect, factors large enough to detect but not large enough to improve the estimate, and two classes of factors large enough to be useful. Much of the advantage of bi-cross-validation comes from cases with factors large enough to detect but too small to be well estimated. We also find that a form of early stopping regularization improves the recovery of the signal matrix.
Recommendations
- Bi-cross-validation of the SVD and the nonnegative matrix factorization
- Determining the number of factors in approximate factor models by twice K-fold cross validation
- Factor Analysis Revisited – How Many Factors are There?
- Determining the number of factors when the number of factors can increase with sample size
- Model selection for factor analysis: some new criteria and performance comparisons
Cites work
- A general framework for multiple testing dependence
- A rationale and test for the number of factors in factor analysis
- A review of signal subspace speech enhancement and its application to noise robust speech recognition
- A Testing Procedure for Determining the Number of Factors in Approximate Factor Models With Large Datasets
- Asymptotic analysis of the squared estimation error in misspecified factor models
- Asymptotics of sample eigenstructure for a large dimensional spiked covariance model
- Asymptotics of the principal components estimator of large factor models with weakly influential factors
- Bi-cross-validation for factor analysis
- Bi-cross-validation of the SVD and the nonnegative matrix factorization
- Boosting as a regularized path to a maximum margin classifier
- Boosting with early stopping: convergence and consistency
- Detection of signals by information theoretic criteria: general asymptotic performance analysis
- Determining the number of components from the matrix of partial correlations
- Determining the Number of Factors in Approximate Factor Models
- Determining the Number of Factors in the General Dynamic Factor Model
- Eigenvalue ratio test for the number of factors
- Eigenvalues of large sample covariance matrices of spiked population models
- Equivalence of regularization and truncated iteration in the solution of ill-posed image reconstruction problems
- Factor modeling for high-dimensional time series: inference for the number of factors
- Finite sample approximation results for principal component analysis: A matrix perturbation approach
- How many principal components? Stopping rules for determining the number of non-trivial axes revisited
- scientific article; zbMATH DE number 3092167 (Why is no real title available?)
- Improved penalization for determining the number of factors in approximate factor models
- Latent variable graphical model selection via convex optimization
- Multiple hypothesis testing adjusted for latent variables, with an application to the AGEMAP gene expression data
- Networks, crowds and markets. Reasoning about a highly connected world.
- On a Heuristic Method of Test Construction and its use in Multivariate Analysis
- On early stopping in gradient descent learning
- OptShrink: An Algorithm for Improved Low-Rank Signal Matrix Denoising by Optimal, Data-Driven Singular Value Shrinkage
- Principal component analysis.
- Sample Eigenvalue Based Detection of High-Dimensional Signals in White Noise Using Relatively Few Samples
- Selecting the number of principal components: estimation of the true rank of a noisy matrix
- Statistical analysis of factor models of high dimension
- TESTS OF SIGNIFICANCE FOR THE LATENT ROOTS OF COVARIANCE AND CORRELATION MATRICES
- The Generalized Dynamic Factor Model
- The Optimal Hard Threshold for Singular Values is <inline-formula> <tex-math notation="TeX">\(4/\sqrt {3}\) </tex-math></inline-formula>
- The singular values and vectors of low rank perturbations of large rectangular random matrices
Cited in
(32)- Bi-cross-validation for factor analysis
- esaBcv
- Bi-cross-validation of the SVD and the nonnegative matrix factorization
- Consistently recovering the signal from noisy functional data
- Prediction in functional regression with discretely observed and noisy covariates
- Preprocessing noisy functional data: a multivariate perspective
- Heteroskedastic PCA: algorithm, optimality, and applications
- Sparse latent factor regression models for genome-wide and epigenome-wide association studies
- Hypothesis tests for principal component analysis when variables are standardized
- Deterministic parallel analysis: an improved method for selecting factors and principal components
- Exploratory bi-factor analysis: the oblique case
- Empirical Bayes matrix factorization
- A Matrix-Free Likelihood Method for Exploratory Factor Analysis of High-Dimensional Gaussian Data
- Structured latent factor analysis for large-scale data: identifiability, estimability, and their implications
- Unifying and generalizing methods for removing unwanted variation based on negative controls
- Improved shrinkage prediction under a spiked covariance structure
- Estimating and Accounting for Unobserved Covariates in High-Dimensional Correlated Data
- Identifying Effects of Multiple Treatments in the Presence of Unmeasured Confounding
- A central limit theorem for the Benjamini-Hochberg false discovery proportion under a factor model
- Safety signal detection with control of latent factors
- Bayesian generalized linear low rank regression models for the detection of vaccine-adverse event associations
- Simultaneous Estimation of Multiple Treatment Effects from Observational Studies
- Entrywise splitting cross-validation in generalized factor models: from sample splitting to entrywise splitting
- An ensemble approach to determine the number of latent dimensions and assess its reliability
- Simultaneous Inference for Generalized Linear Models with Unmeasured Confounders
- Robust Estimation for Number of Factors in High Dimensional Factor Modeling via Spearman Correlation Matrix
- Sparse Bayesian factor analysis when the number of factors is unknown (with discussion)
- High-dimensional large-scale mixed-type data imputation under missing at random
- Data Thinning for Poisson Factor Models and its Applications
- Mediation analysis with unmeasured confounding between parallel mediators and outcome
- Sparse generalized factor models with weaker loadings
- A tutorial on Bayesian multi-study factor analysis with applications in nutrition and genomics
This page was built for publication: Bi-cross-validation for factor analysis
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q104117)