Nonparametric Identifiability of Causal Representations from Unknown Interventions
Julius von Kügelgen Max Planck Institute for Intelligent Systems, Tübingen, Germany Department of Engineering, University of Cambridge, United Kingdom Michel Besserve Max Planck Institute for Intelligent Systems, Tübingen, Germany Liang Wendong Max Planck Institute for Intelligent Systems, Tübingen, Germany Luigi Gresele Max Planck Institute for Intelligent Systems, Tübingen, Germany Armin Kekić Max Planck Institute for Intelligent Systems, Tübingen, Germany Elias Bareinboim Columbia University, USA David M. Blei Columbia University, USA Bernhard Schölkopf Max Planck Institute for Intelligent Systems, Tübingen, Germany
Abstract
We study causal representation learning, the task of inferring latent causal variables and their causal relations from high-dimensional functions (“mixtures”) of the variables. Prior work relies on weak supervision, in the form of counterfactual pre- and post-intervention views or temporal structure; places restrictive assumptions, such as linearity, on the mixing function or latent causal model; or requires partial knowledge of the generative process, such as the causal graph or intervention targets. We instead consider the general setting in which both the causal model and the mixing function are nonparametric. The learning signal takes the form of multiple datasets, or environments, arising from unknown interventions in the underlying causal model. Our goal is to identify both the ground truth latents and their causal graph up to a set of ambiguities which we show to be irresolvable from interventional data. We study the fundamental setting of two causal variables and prove that the observational distribution and one perfect intervention per node suffice for identifiability, subject to a genericity condition. This condition rules out spurious solutions that involve fine-tuning o
中文速览
从高维观测数据(比如图像或传感器信号)中还原出背后真正起因果作用的潜变量及其因果关系图,是因果表示学习(causal representation learning)的核心难题,而在混合函数和因果机制都完全未知、非参数的最一般情形下,这个问题此前几乎没有可识别性保证。本文证明:只要能访问来自同一因果模型的多个"干预环境"(interventional environments),对于两个潜变量的情形,一个观测分布加上对每个节点各做一次完美干预就足以识别潜变量及其因果图;对任意数量变量,每个节点有两个不同的完美干预环境即可保证可识别性,且所有等价解都保留了潜变量之间因果影响的强度信息。实验进一步表明,只有采用正确因果结构的生成模型才能同时达到最佳拟合并还原真实潜变量。这项工作首次为完全非参数、干预未知的因果表示学习建立了严格的可识别性理论,划清了在缺乏更多监督信号时该任务的能力边界。
原文 arXiv:2306.00542;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2306.00542v2