When Do Neural Networks Outperform Kernel Methods?
Behrooz Ghorbani Thanks: Department of Electrical Engineering, Stanford University Song Mei Thanks: Department of Statistics, University of California, Berkeley Theodor Misiakiewicz Thanks: Department of Statistics, Stanford University Andrea Montanari Thanks: Google Research, Brain Team
Abstract
For a certain scaling of the initialization of stochastic gradient descent (SGD), wide neural networks (NN) have been shown to be well approximated by reproducing kernel Hilbert space (RKHS) methods. Recent empirical work showed that, for some classification tasks, RKHS methods can replace NNs without a large loss in performance. On the other hand, two-layers NNs are known to encode richer smoothness classes than RKHS and we know of special examples for which SGD-trained NN provably outperform RKHS. This is true even in the wide network limit, for a different scaling of the initialization.
原文 arXiv:2006.13409;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2006.13409v2