Positively Scale-Invariant Flatness of ReLU Neural Networks
Mingyang Yi Academy of Mathematics and Systems Science, Chinese Academy of Sciences University of Chinese Academy of Sciences Qi Meng Microsoft Research Wei Chen Microsoft Research Zhi-ming Ma Academy of Mathematics and Systems Science, Chinese Academy of Sciences University of Chinese Academy of Sciences Tie-Yan Liu Microsoft Research
Abstract
It was empirically confirmed by Keskar et al.[10] that flatter minima generalize better. However, for the popular ReLU network, sharp minimum can also generalize well [3]. The conclusion demonstrates that the existing definitions of flatness fail to account for the complex geometry of ReLU neural networks because they can’t cover the Positively Scale-Invariant (PSI) property of ReLU network. In this paper, we formalize the PSI causes problem of existing definitions of flatness and propose a new description of flatness - PSI-flatness. PSI-flatness is defined on the values of basis paths [16] instead of weights. Values of basis paths have been shown to be the PSI-variables and can sufficiently represent the ReLU neural networks which ensure the PSI property of PSI-flatness. Then we study the relation between PSI-flatness and generalization theoretically and empirically. First, we formulate a generalization bound based on PSI-flatness which shows generalization error decreasing with the ratio between the largest basis path value and the smallest basis path value. That is to say, the minimum with balanced values of basis paths will more likely to be flatter and generalize better. Final
中文速览
深度神经网络(DNN)在实践中泛化能力出奇地好,研究者普遍认为"平坦极小值"是关键——损失曲面越平坦,模型越不容易过拟合。然而,现有的平坦度定义都建立在网络权重空间上,而ReLU网络天生具有"正尺度不变性"(Positively Scale-Invariant, PSI):把某一隐藏节点的入边权重乘以正常数c、出边权重除以c,网络输出完全不变,这意味着同一个函数对应无穷多组权重,基于权重的平坦度度量会因此失真,有人甚至能构造出"泛化好但按现有定义极其不平坦"的反例。本文的核心贡献是提出PSI-flatness:把平坦度改定义在"基路径值"(basis path values)空间上——基路径值天然是PSI不变量,能唯一表征ReLU网络的计算函数——从而绕开了尺度变换带来的歧义;理论上,作者推导出一个PAC-Bayes泛化误差界,表明基路径值之间越均衡(最大值与最小值之比越小),PSI-flatness越小,泛化误差越低;实验上,可视化损失曲面也印证了PSI-flatness更小的极小值点确实处于更"平坦"的谷底。这项工作从几何角度为理解ReLU网络的泛化提供了更坚实的理论基础,也为设计鼓励均衡参数分布的训练算法提供了新动机。
原文 arXiv:1903.02237;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1903.02237v1