What Can ResNet Learn Efficiently, Going Beyond Kernels?Thanks: V1 appears on this date, V2 slightly improved the lower bound, V3 strengthens experiments and adds citation to “backward feature correction” which is an even stronger form of hierarchical learning [2]. We would like to thank Greg Yang for many enlightening conversations as well as discussions on neural tangent kernels. A 45-min presentation of this result at the UC Berkeley Simons Institute can be found at https://youtu.be/NNPCk2gvTnI.
Zeyuan Allen-Zhu Email: Affiliation: Microsoft Research AI Yuanzhi Li Email: Affiliation: Stanford University
Abstract
How can neural networks such as ResNet efficiently learn CIFAR-10 with test accuracy more than $96\%$ , while other methods, especially kernel methods, fall relatively behind? Can we more provide theoretical justifications for this gap?
原文 arXiv:1905.10337;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1905.10337v3