What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation
Vitaly Feldman Thanks: Equal contribution. Thanks: Part of this work done while the author was at Google Research, Brain Team. Affiliation: Apple Chiyuan Zhang Affiliation: Google Research, Brain Team
Abstract
Deep learning algorithms are well-known to have a propensity for fitting the training data very well and often fit even outliers and mislabeled data points. Such fitting requires memorization of training data labels, a phenomenon that has attracted significant research interest but has not been given a compelling explanation so far. A recent work of [Fel19] proposes a theoretical explanation for this phenomenon based on a combination of two insights. First, natural image and data distributions are (informally) known to be long-tailed, that is have a significant fraction of rare and atypical examples. Second, in a simple theoretical model such memorization is necessary for achieving close-to-optimal generalization error when the data distribution is long-tailed. However, no direct empirical evidence for this explanation or even an approach for obtaining such evidence were given.
原文 arXiv:2008.03703;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2008.03703v1