Reconciling modern machine learning practice and the bias-variance trade-off
Mikhail Belkin Affiliation: The Ohio State University, Columbus, OH Daniel Hsu Affiliation: Columbia University, New York, NY Siyuan Ma Affiliation: The Ohio State University, Columbus, OH Soumik Mandal Affiliation: The Ohio State University, Columbus, OH
Abstract
Breakthroughs in machine learning are rapidly changing science and society, yet our fundamental understanding of this technology has lagged far behind. Indeed, one of the central tenets of the field, the bias-variance trade-off, appears to be at odds with the observed behavior of methods used in the modern machine learning practice. The bias-variance trade-off implies that a model should balance under-fitting and over-fitting: rich enough to express underlying structure in data, simple enough to avoid fitting spurious patterns. However, in the modern practice, very rich models such as neural networks are trained to exactly fit (i.e., interpolate) the data. Classically, such models would be considered over-fit, and yet they often obtain high accuracy on test data. This apparent contradiction has raised questions about the mathematical foundations of machine learning and their relevance to practitioners.
原文 arXiv:1812.11118;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1812.11118v2