A Convergence Theory for Deep Learning via Over-Parameterization Thanks: This paper was presented at ICML 2019, and is a simplified version of our prior work for recurrent neural networks (RNNs) [6]. V1 appears on arXiv on this date and no new result is added since then. V2 adds citations, V3/V4/V5 polish writing. This work was done when Yuanzhi Li and Zhao Song were 2018 summer interns at Microsoft Research Redmond. When this work was performed, Yuanzhi Li was also affiliated with Princeton University, and Zhao Song was also affiliated with UW and Harvard. We would like to specially thank Greg Yang for many enlightening discussions, thank Ofer Dekel, Sebastien Bubeck, and Harry Shum for very helpful conversations, and thank Jincheng Mei for carefully checking the proofs of this paper.
Zeyuan Allen-Zhu Email: Affiliation: Microsoft Research AI Yuanzhi Li Email: Affiliation: Stanford University Affiliation: Princeton University Zhao Song Email: Affiliation: UT-Austin Affiliation: University of Washington Affiliation: Harvard University
Abstract
Deep neural networks (DNNs) have demonstrated dominating performance in many fields; since AlexNet, networks used in practice are going wider and deeper. On the theoretical side, a long line of works has been focusing on training neural networks with one hidden layer. The theory of multi-layer networks remains largely unsettled.
原文 arXiv:1811.03962;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1811.03962v5