Depth-Width Trade-offs for Neural Networks via Topological Entropy
Kaifeng Bu1 Email address: (K.Bu) , Yaobo Zhang2,3 Email address: (Y.Zhang) and Qingxian Luo4,5 Email address: $1$Department of Physics, Harvard University, Cambridge, Massachusetts 02138, USA $2$Zhejiang Institute of Modern Physics, Zhejiang University, Hangzhou, Zhejiang 310027, China $3$Department of Physics, Zhejiang University, Hangzhou Zhejiang 310027, China $4$School of Mathematical Sciences, Zhejiang University, Hangzhou, Zhejiang 310027, China $5$Center for Data Science, Zhejiang University, Hangzhou Zhejiang 310027, China
Abstract
One of the central problems in the study of deep learning theory is to understand how the structure properties, such as depth, width and the number of nodes, affect the expressivity of deep neural networks. In this work, we show a new connection between the expressivity of deep neural networks and topological entropy from dynamical system, which can be used to characterize depth-width trade-offs of neural networks. We provide an upper bound on the topological entropy of neural networks with continuous semi-algebraic units by the structure parameters. Specifically, the topological entropy of ReLU network with $l$ layers and $m$ nodes per layer is upper bounded by $O(l\log m)$ . Besides, if the neural network is a good approximation of some function $f$ , then the size of the neural network has an exponential lower bound with respect to the topological entropy of $f$ . Moreover, we discuss the relationship between topological entropy, the number of oscillations, periods and Lipschitz constant.
原文 arXiv:2010.07587;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2010.07587v1