Reconstructing a neural net from its output
Summary: Neural nets were originally introduced as highly simplified models of the nervous system. Today they are widely used in technology and studied theoretically by scientists from several disciplines. However, they remain little understood. Mathematically, a (feed-forward) neural net consists of (1) a finite sequence of positive integers \((D_0,D_1,\dots,D_L)\), (2) a family of real numbers \((\omega^\ell_{jk})\) defined for \(1\leq\ell\leq L\), \(1\leq j\leq D_\ell\), \(1\leq k\leq D_{\ell-1}\), and (3) a family of real numbers \((\theta^\ell_j)\) defined for \(1\leq\ell\leq L\), \(1\leq j\leq D_\ell\). The sequence \((D_0,D_1,\dots,D_L)\) is called the architecture of the neural net, while the \(\omega^\ell_{jk}\) are called weights and the \(\theta^\ell_j\) thresholds. Neural nets are used to compute nonlinear maps from \(\mathbb{R}^N\) to \(\mathbb{R}^M\) by the following construction. We begin by fixing a nonlinear function \(\sigma(x)\) of one variable. Analogy with the nervous system suggests that we take \(\sigma(t)\) asymptotic to constants as \(t\) tends to \(\pm\infty\); a standard choice, which we adopt throughout this paper, is \(\sigma(x)=\text{tanh}(x/2)\). Given an ``input \((t_1,\dots,t_{D_0})\in\mathbb{R}^{D_0}\), we define real numbers \(x^\ell_j\) for \(0\leq\ell\leq L\), \(1\leq j\leq D_\ell\) by the following induction on \(\ell\). If \(\ell=0\) then \(x^\ell_j= t_j\). If the \(x^{\ell-1}_k\) are known, \(1\leq\ell\leq L\), then we set \[ x^\ell_j=\sigma\Biggl( \sum_{1\leq k\leq D_{\ell-1}}\omega^\ell_{jk} x^{\ell-1}_k+ \theta^\ell_j\Biggr),\quad\text{for } 1\leq j\leq D_\ell. \] Here \(x^\ell_1,\dots, x^\ell_{D_\ell}\) are interpreted as the output of \(D_\ell\) ``neurons in the \(\ell\)th ``layer of the net. The output map of the net is defined as the map \[ \Phi: (t_1,\dots, t_{D_0})\mapsto (x^L_1,\dots, x^L_{D_L}). \] In practical applications, one tries to pick the neural net \[ [(D_0,D_1,\dots, D_L),\;(\omega^\ell_{jk}),\;(\theta^\ell_j)] \] so that the output map \(\Phi\) approximates a given map about which we have only imperfect information. The main result of this paper is that under generic conditions, perfect knowledge of the output map \(\Phi\) uniquely specifies the architecture, the weights and the thresholds of a neural net, up to obvious symmetries.
- On the identification of neural responses
- Identifying linear combinations of ridge functions
- Ehresmann connections and feedforward neural networks
- Neural networks as set-valued dynamical systems and the universality of the windowed Fourier transform
- Neural networks, rational functions, and realization theory
- Stable recovery of entangled weights: towards robust identification of deep neural networks from minimal samples
- Information theory and recovery algorithms for data fusion in Earth observation
- Neural network identifiability for a family of sigmoidal nonlinearities
- Robust and resource-efficient identification of two hidden layer neural networks
- Metric entropy limits on recurrent neural network learning of linear dynamical systems
- Affine symmetries and neural network identifiability
- Absence of bottlenecks in a neural network determines its generic functional properties
- An embedding of ReLU networks and an analysis of their identifiability
- A mathematical solution to a network construction problem.
- scientific article; zbMATH DE number 2089826 (Why is no real title available?)
- How to modify a neural network gradually without changing its input-output functionality
- scientific article; zbMATH DE number 2044721 (Why is no real title available?)
- scientific article; zbMATH DE number 1528653 (Why is no real title available?)
- Construction of neural networks for determination of singular points of linear complexes of planes of category B
- Higher-order quasi-Monte Carlo training of deep neural networks
- Using neural nets for eliminating noise in experimental signals
- Harmonic functions on the closed cube: an application to learning theory
- Parameter identifiability of a deep feedforward ReLU neural network
- Polynomial time cryptanalytic extraction of neural network models
- On partial groupoids associated with the composition of multilayer feedforward neural networks
- Extracting some layers of deep neural networks in the hard-label setting
- Polynomial time cryptanalytic extraction of deep neural networks in the hard-label setting
- Hard-label cryptanalytic extraction of neural network models
- Efficient identification of wide shallow neural networks with biases
- Navigating the deep: end-to-end extraction on deep neural networks
This page was built for publication: Reconstructing a neural net from its output
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1344572)