Exponential Family Estimation via Adversarial Dynamics Embedding
∗Bo Dai1 Zhen Liu2 indicates equal contribution. Email: {bodai, ∗Hanjun Dai1 Niao He3 Arthur Gretton4 Le Song5,6 Dale Schuurmans1,7 1Google Research Brain Team 2Mila University of Montreal 3University of Illinois at Urbana Champaign 4University College London 5Georgia Institute of Technology 6Ant Financial 7University of Alberta
Abstract
We present an efficient algorithm for maximum likelihood estimation (MLE) of exponential family models, with a general parametrization of the energy function that includes neural networks. We exploit the primal-dual view of the MLE with a kinetics augmented model to obtain an estimate associated with an adversarial dual sampler. To represent this sampler, we introduce a novel neural architecture, dynamics embedding, that generalizes Hamiltonian Monte-Carlo (HMC). The proposed approach inherits the flexibility of HMC while enabling tractable entropy estimation for the augmented model. By learning both a dual sampler and the primal model simultaneously, and sharing parameters between them, we obviate the requirement to design a separate sampling procedure once the model has been trained, leading to more effective learning. We show that many existing estimators, such as contrastive divergence, pseudo/composite-likelihood, score matching, minimum Stein discrepancy estimator, non-local contrastive objectives, noise-contrastive estimation, and minimum probability flow, are special cases of the proposed approach, each expressed by a different (fixed) dual sampler. An empirical investigati
中文速览
最大似然估计(MLE)是拟合指数族模型(含神经网络参数化的能量模型)的黄金标准,但配分函数不可解析计算让它在实践中极难使用。这篇论文提出了一种叫做"对抗性动力学嵌入"(Adversarial Dynamics Embedding,ADE)的新算法:借助原始-对偶分解把 MLE 改写成一个极大极小问题,再用一种仿照哈密顿蒙特卡洛(HMC)的神经网络架构同时表示"采样器"(对偶分布)和"模型"(原始分布),两者共享参数、联合训练,从而避免了单独设计采样程序的麻烦,并保证熵项可解析计算。作者还证明,对比散度、伪似然、得分匹配、噪声对比估计等一系列经典方法都是 ADE 在固定采样器时的特例。实验表明,让采样器随模型一起自适应学习能显著超越现有最优估计器,为训练指数族和能量模型提供了一条更高效、更通用的路径。
原文 arXiv:1904.12083;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1904.12083v3