The Description Length of Deep Learning Models
Léonard Blier École Normale Supérieure Paris, France、Yann Ollivier Facebook Artificial Intelligence Research Paris, France
Abstract
Solomonoff’s general theory of inference (Solomonoff, 1964) and the Minimum Description Length principle (Grünwald, 2007, Rissanen, 2007) formalize Occam’s razor, and hold that a good model of data is a model that is good at losslessly compressing the data, including the cost of describing the model itself. Deep neural networks might seem to go against this principle given the large number of parameters to be encoded.
原文 arXiv:1802.07044;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1802.07044v5