Prevalence of neural collapse during the terminal phase of deep learning training
From MaRDI portal
Publication:5073172
Abstract: Modern practice for training classification deepnets involves a Terminal Phase of Training (TPT), which begins at the epoch where training error first vanishes; During TPT, the training error stays effectively zero while training loss is pushed towards zero. Direct measurements of TPT, for three prototypical deepnet architectures and across seven canonical classification datasets, expose a pervasive inductive bias we call Neural Collapse, involving four deeply interconnected phenomena: (NC1) Cross-example within-class variability of last-layer training activations collapses to zero, as the individual activations themselves collapse to their class-means; (NC2) The class-means collapse to the vertices of a Simplex Equiangular Tight Frame (ETF); (NC3) Up to rescaling, the last-layer classifiers collapse to the class-means, or in other words to the Simplex ETF, i.e. to a self-dual configuration; (NC4) For a given activation, the classifier's decision collapses to simply choosing whichever class has the closest train class-mean, i.e. the Nearest Class Center (NCC) decision rule. The symmetric and very simple geometry induced by the TPT confers important benefits, including better generalization performance, better robustness, and better interpretability.
Cites work
- A Mathematical Theory of Deep Convolutional Neural Networks for Feature Extraction
- Adversarial noise attacks of deep learning architectures: stability analysis via sparse-modeled signals
- Convolutional neural networks analyzed via convolutional sparse coding
- Grassmannian frames with applications to coding and communication
- Group invariant scattering
- scientific article; zbMATH DE number 1158743 (Why is no real title available?)
- Multi-layer sparse coding: the holistic way
- Reconciling modern machine-learning practice and the classical bias-variance trade-off
- The implicit bias of gradient descent on separable data
Cited in
(18)- Deep regularization and direct training of the inner layers of neural networks with kernel flows
- Neural collapse under cross-entropy loss
- Neural collapse with unconstrained features
- Learning sparse features can lead to overfitting in neural networks
- Theoretical guarantees for low-rank compression of deep neural networks
- Gradient flow in parameter space is equivalent to linear interpolation in output space
- A deep neural network framework for multivalued mapping problems with varying cardinality and its applications to imaging
- The mathematics of adversarial attacks in AI -- why deep learning is unstable despite the existence of stable neural networks
- Interpretable global minima of deep ReLU neural networks on sequentially separable data
- Spectral alignment of stochastic gradient descent for high-dimensional classification tasks
- On non-approximability of zero loss global \(\mathcal{L}^2\) minimizers by gradient descent in deep learning
- Semi-Supervised Triply Robust Inductive Transfer Learning
- Modern and emerging phenomena in machine learning. Abstracts from the workshop held March 8--13, 2026
- Understanding deep representation learning via layerwise feature compression and discrimination
- Stochastic gradient descent in high dimensions for multi-spiked tensor PCA
- SAMix: calibrated and accurate continual learning via sphere-adaptive mixup and neural collapse
- Local geometry of high-dimensional mixture models: effective spectral theory and dynamical transitions
- Singular parameters and missing limits in neural PDE solvers
This page was built for publication: Prevalence of neural collapse during the terminal phase of deep learning training
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5073172)