I Don’t Need 𝐮𝐮\mathbf{u}: Identifiable Non-Linear ICA Without Side Information
\nameMatthew Willetts \addrUniversity College London、The Alan Turing Institute \AND\nameBrooks Paige \addrUniversity College London、The Alan Turing Institute
Abstract
In this paper, we investigate the algorithmic stability of unsupervised representation learning with deep generative models, as a function of repeated re-training on the same input data. Algorithms for learning low dimensional linear representations—for example principal components analysis (PCA), or linear independent components analysis (ICA)—come with guarantees that they will always reveal the same latent representations (perhaps up to an arbitrary rotation or permutation). Unfortunately, for non-linear representation learning, such as in a variational auto-encoder (VAE) model trained by stochastic gradient descent, we have no such guarantees. Recent work on identifiability in non-linear ICA have introduced a family of deep generative models that have identifiable latent representations, achieved by conditioning on side information (e.g. informative labels). We empirically evaluate the stability of these models under repeated re-estimation of parameters, and compare them to both standard VAEs and deep generative models which learn to cluster in their latent space. Surprisingly, we discover side information is not necessary for algorithmic stability: using standard quantitative
中文速览
深度生成模型(deep generative model, DGM)每次用随机梯度下降重新训练,学到的潜变量表示会不会每次都一样,这是一个关乎模型是否真正"有意义"的核心问题。现有理论表明,给模型加入辅助侧信息(如类别标签)可以让非线性独立成分分析(non-linear ICA)的潜变量表示具有可识别性(identifiability),即不同训练结果最多只相差一个仿射变换;但不依赖侧信息的纯无监督模型理论上没有这种保证。本文系统比较了标准VAE、在潜空间做聚类的VaDE,以及依赖标签侧信息的可识别VAE(iVAE)在反复重新训练时表示的一致性,发现一个出人意料的结果:VaDE在没有任何侧信息的情况下,其表示的可识别性与有理论保证的iVAE相当甚至更好,两者之间的差异在统计上不显著。这一发现挑战了"必须依赖侧信息才能获得稳定可识别表示"的主流理论假设,也对当前可识别性领域的实验评估方法提出了质疑,为理解无监督深度生成模型的可靠性提供了新的实证依据。
原文 arXiv:2106.05238;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2106.05238v4