Gaussian Process Behaviour in Wide Deep Neural Networks
Alexander G. de G. Matthews Affiliation: Department of Engineering Affiliation: Trumpington Street Affiliation: University of Cambridge, UK Mark Rowland Affiliation: Department of Pure Mathematics and Mathematical Statistics Affiliation: Wilberforce Road Affiliation: University of Cambridge, UK Jiri Hron Affiliation: Department of Engineering Affiliation: Trumpington Street Affiliation: University of Cambridge, UK Richard E. Turner Affiliation: Department of Engineering Affiliation: Trumpington Street Affiliation: University of Cambridge, UK Zoubin Ghahramani Affiliation: Department of Engineering Affiliation: Trumpington Street Affiliation: University of Cambridge, UK Affiliation: Uber AI Labs
Abstract
Whilst deep neural networks have shown great empirical success, there is still much work to be done to understand their theoretical properties. In this paper, we study the relationship between random, wide, fully connected, feedforward networks with more than one hidden layer and Gaussian processes with a recursive kernel definition. We show that, under broad conditions, as we make the architecture increasingly wide, the implied random function converges in distribution to a Gaussian process, formalising and extending existing results by Neal 1996 to deep networks. To evaluate convergence rates empirically, we use maximum mean discrepancy. We then compare finite Bayesian deep networks from the literature to Gaussian processes in terms of the key predictive quantities of interest, finding that in some cases the agreement can be very close. We discuss the desirability of Gaussian process behaviour and review non-Gaussian alternative models from the literature.11 1 Code for the experiments in the paper can be found at https://github.com/widedeepnetworks/widedeepnetworks
原文 arXiv:1804.11271;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1804.11271v2