Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot Affiliation: École Polytechnique Fédérale de Lausanne Email: Franck Gabriel Affiliation: Imperial College London and École Polytechnique Fédérale de Lausanne Email: Clément Hongler Affiliation: École Polytechnique Fédérale de Lausanne Email:
Abstract
At initialization, artificial neural networks (ANNs) are equivalent to Gaussian processes in the infinite-width limit (16; 4; 7; 13; 6), thus connecting them to kernel methods. We prove that the evolution of an ANN during training can also be described by a kernel: during gradient descent on the parameters of an ANN, the network function $f_{\theta}$ (which maps input vectors to output vectors) follows the kernel gradient of the functional cost (which is convex, in contrast to the parameter cost) w.r.t. a new kernel: the Neural Tangent Kernel (NTK). This kernel is central to describe the generalization features of ANNs. While the NTK is random at initialization and varies during training, in the infinite-width limit it converges to an explicit limiting kernel and it stays constant during training. This makes it possible to study the training of ANNs in function space instead of parameter space. Convergence of the training can then be related to the positive-definiteness of the limiting NTK. We prove the positive-definiteness of the limiting NTK when the data is supported on the sphere and the non-linearity is non-polynomial.
原文 arXiv:1806.07572;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1806.07572v4