Geometry of linear neural networks: equivariance and invariance under permutation groups (Q6972316)
From MaRDI portal
!
This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use the normal view instead:
scientific article; zbMATH DE number 8052175
| Language | Label | Description | Also known as |
|---|---|---|---|
| default for all languages | No label defined |
||
| English | Geometry of linear neural networks: equivariance and invariance under permutation groups |
scientific article; zbMATH DE number 8052175 |
Statements
Geometry of linear neural networks: equivariance and invariance under permutation groups (English)
0 references
12 June 2025
0 references
Understanding the geometry of the function space of a neural network architecture is a fundamental task in the theory of deep learning. Here, the authors specialize to the case of linear networks -- neural networks built as a composition of linear functions -- to study invariance and equivariance properties under a group action.\N\NFormally, a linear neural network represents a linear function of bounded rank, and therefore gives a point on the determinantal variety \(\mathcal{M}_{r,m\times n}\) of \(m \times n\) matrices with rank at most \(r\). If \(G \leq S_n\) is a subgroup of the permutation group, then the set of \(G\)-invariant linear functions \(\mathcal{I}_{r,m\times n}^G\) is an irreducible subvariety of \(\mathcal{M}_{r,m\times n}\) and the authors compute its dimension, degree, and singular locus. The case of equivariant functions is more subtle, and the authors describe the irreducible components of the set of \(G\)-equivariant functions \(\mathcal{E}^G_{r,n \times n}\) for cyclic subgroups \(G = \langle \sigma \rangle \leq S_n\). These irreducible components are indexed by integer vectors \(\mathbf{r}\) corresponding to certain partitions of the rank \(r\) (Theorem 4.2). For each irreducible component \(\mathcal{E}^{G,\mathbf{r}}_{r,n\times n}(\mathbb{C}) \subseteq \mathcal{E}^{G}_{r,n\times n}(\mathbb{C})\), the authors compute the dimension, degree, and singular locus, as well as the real dimension and singular locus of \(\mathcal{E}^{G,\mathbf{r}}_{r,n\times n}(\mathbb{R})\). As a final consideration, the authors compute the squared error degree of these subvarieties of \(\mathcal{M}_{r,m\times n}\), defined as the number of complex critical points on the variety of the function \(M \mapsto \|MX - Y\|^2_F\) for generic data matrices \(X\) and \(Y\). The authors connect these results to the design of neural networks by describing weight sharing strategies which yield networks describing \(\mathcal{I}^G_{r,m\times n}\) or an irreducible component of \(\mathcal{E}^{G}_{r, n \times n}\).\N\NThe theory developed by the authors is accompanied by helpful small-scale examples as well as numerical experiments. The numerical experiments provide evidence that the substantial computational benefit obtained by imposing equivariance on the network architecture does not lead to a substantial loss in model accuracy.
0 references
neuromanifold
0 references
linear neural network
0 references
equivariance
0 references
determinantal variety
0 references
squared-error loss
0 references
Eckart-Young
0 references