Overparameterization without Overfitting: Jacobian-based Generalization Guarantees for Neural Networks
Samet Oymak Zalan Fabian α Mingchen Li α Mahdi Soltanolkotabi Department of Electrical and Computer Engineering, University of California, Riverside, CAMing Hsieh Department of Electrical Engineering, University of Southern California, Los Angeles, CADepartment of Computer Science and Engineering, University of California, Riverside, CA
Abstract
Modern neural network architectures often generalize well despite containing many more parameters than the size of the training dataset. This paper explores the generalization capabilities of neural networks trained via gradient descent. We develop a data-dependent optimization and generalization theory which leverages the low-rank structure of the Jacobian matrix associated with the network. Our results help demystify why training and generalization is easier on clean and structured datasets and harder on noisy and unstructured datasets as well as how the network size affects the evolution of the train and test errors during training. Specifically, we use a control knob to split the Jacobian spectum into “information" and “nuisance" spaces associated with the large and small singular values. We show that over the information space learning is fast and one can quickly train a model with zero training loss that can also generalize well. Over the nuisance space training is slower and early stopping can help with generalization at the expense of some bias. We also show that the overall generalization capability of the network is controlled by how well the label vector is aligned with
原文 arXiv:1906.05392;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1906.05392v2