Better Mixing via Deep Representations
Yoshua Bengio, Grégoire Mesnil, Yann Dauphin and Salah Rifai Department of Computer Science and Operations Research University of Montreal Montreal, H3C 3J7
Abstract
It has previously been hypothesized, and supported with some experimental evidence, that deeper representations, when well trained, tend to do a better job at disentangling the underlying factors of variation. We study the following related conjecture: better representations, in the sense of better disentangling, can be exploited to produce faster-mixing Markov chains. Consequently, mixing would be more efficient at higher levels of representation. To better understand why and how this is happening, we propose a secondary conjecture: the higher-level samples fill more uniformly the space they occupy and the high-density manifolds tend to unfold when represented at higher levels. The paper discusses these hypotheses and tests them experimentally through visualization and measurements of mixing and interpolating between samples.
中文速览
深度神经网络越深、表征越抽象,生成样本时的马尔可夫链(Markov chain)混合速度是否也越快?这篇论文围绕这一核心猜想展开,提出并检验了三个层层递进的假设:更深的表征能更好地"解耦"数据的内在变化因素,这种解耦使高密度流形(manifold)在表征空间中趋于展开并更均匀地铺满所占空间,从而让采样链更容易在不同模式之间跳转。作者用深度置信网络(Deep Belief Network,DBN)和收缩自编码器(Contractive Auto-Encoder,CAE)在人脸和手写数字数据上做了实验,通过可视化、混合速度测量和样本插值等手段,发现更高层的表征空间确实呈现出更均匀的分布和更平滑的流形结构,使得链的混合效率显著提升。这一发现有助于从根本上解释为何深层生成模型能产生更高质量的样本,也为理解深度表征的本质优势提供了新的视角。
原文 arXiv:1207.4404;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1207.4404v1