Interpretable global minima of deep ReLU neural networks on sequentially separable data (Q6887366)
From MaRDI portal
!
This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use the normal view instead:
scientific article; zbMATH DE number 8133728
| Language | Label | Description | Also known as |
|---|---|---|---|
| default for all languages | No label defined |
||
| English | Interpretable global minima of deep ReLU neural networks on sequentially separable data |
scientific article; zbMATH DE number 8133728 |
Statements
Interpretable global minima of deep ReLU neural networks on sequentially separable data (English)
0 references
9 December 2025
0 references
This interesting paper studies interpretable global minima of deep ReLU neural networks on sequentially separable data. More specifically, the paper gives an interpretation for the action of the layers of a neural network, and explicitly constructs global minima for a classification task with well distributed data, in a way that is independent of the number of training samples. The characterization is achieved via a map encoding the action layer given by a weight matrix say \(W\), a bias vector say \(b\), and activation function say \(\sigma\) in the form of: \((W*)(\sigma(Wx+b)-b)\) with \(W*\) the generalized inverse of \(W\).The configurations for the training data considered are sufficiently small and are well separated clusters corresponding to each class. Morever, equivalence classes are sequentially linearly separable. For \(Q\) classes of data in \(\mathbb R^n\), global minimizers can be described with \(Q(n + 2)\) parameters.\N\NThe paper is well written with a good set of references.
0 references
deep learning
0 references
neural network
0 references
classification
0 references
cluster
0 references
global minima
0 references
data
0 references
0 references