Deep Convolutional Networks as shallow Gaussian Processes
Adrià Garriga-Alonso Affiliation: department of Engineering Affiliation: University of Cambridge Email: Carl Edward Rasmussen Affiliation: Department of Engineering Affiliation: University of Cambridge Email: Laurence Aitchison Affiliation: Department of Engineering Affiliation: University of Cambridge Email:
Abstract
We show that the output of a (residual) convolutional neural network with an appropriate prior over the weights and biases is a Gaussian process in the limit of infinitely many convolutional filters, extending similar results for dense networks. For a convolutional neural network, the equivalent kernel can be computed exactly and, unlike ‘‘deep kernels’’, has very few parameters: only the hyperparameters of the original convolutional neural network. Further, we show that this kernel has two properties that allow it to be computed efficiently; the cost of evaluating the kernel for a pair of images is similar to a single forward pass through the original convolutional neural network with only one filter per layer. The kernel equivalent to a 32-layer ResNet obtains 0.84% classification error on MNIST, a new record for Gaussian Processes with a comparable number of parameters. 11 1 Code to replicate this paper is available at https://github.com/convnets-as-gps/convnets-as-gps
原文 arXiv:1808.05587;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1808.05587v2