Deep Limits of Residual Neural Networks
Matthew Thorpe Email: Department of Mathematics, University of Manchester, Manchester, M13 9PL The Alan Turing Institute, London, NW1 2DB, UK Yves van Gennip Delft Institute of Applied Mathematics, Delft University of Technology, 2628 CD Delft, The Netherlands
Abstract
Neural networks have been very successful in many applications; we often, however, lack a theoretical understanding of what the neural networks are actually learning. This problem emerges when trying to generalise to new data sets. The contribution of this paper is to show that, for the residual neural network model, the deep layer limit coincides with a parameter estimation problem for a nonlinear ordinary differential equation. In particular, whilst it is known that the residual neural network model is a discretisation of an ordinary differential equation, we show convergence in a variational sense. This implies that optimal parameters converge in the deep layer limit. This is a stronger statement than saying for a fixed parameter the residual neural network model converges (the latter does not in general imply the former). Our variational analysis provides a discrete-to-continuum $\Gamma$ -convergence result for the objective function of the residual neural network training step to a variational problem constrained by a system of ordinary differential equations; this rigorously connects the discrete setting to a continuum problem.
中文速览
残差神经网络(ResNet)虽然在分类、图像等任务上表现出色,但人们对它"究竟在学什么"缺乏严格的数学理解,尤其是当网络层数趋于无穷时,训练出的参数会收敛到什么。这篇论文从变分分析的角度出发,证明了当 ResNet 的层数 $n\to\infty$ 时,网络的训练目标函数(离散的最小化问题)在 $\Gamma$-收敛(Gamma-convergence)意义下收敛到一个由常微分方程(ODE)约束的连续变分问题,从而严格保证了最优参数本身也会收敛——而不仅仅是对固定参数、离散解趋近于 ODE 的解。这一结果将离散的网络训练与连续动力系统紧密地联系起来,不仅为理解 ResNet 的泛化行为提供了理论基础,也为利用庞特里亚金最大值原理等偏微分方程工具来直接求解训练问题、或设计新型网络架构开辟了新路径。
原文 arXiv:1810.11741;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1810.11741v4