Density Estimation in Infinite Dimensional Exponential Families
\nameBharath Sriperumbudur \addrDepartment of Statistics, Pennsylvania State University University Park, PA 16802, USA. \AND\nameKenji Fukumizu \addrThe Institute of Statistical Mathematics 10-3 Midoricho, Tachikawa, Tokyo 190-8562 Japan. \AND\nameArthur Gretton \addrGatsby Computational Neuroscience Unit, University College London Sainsbury Wellcome Centre, 25 Howland Street, London W1T 4JG, UK \AND\nameAapo Hyvärinen \addrDepartment of Computer Science, University of Helsinki P.O. Box 68, FIN-00014, Finland. \AND\nameRevant Kumar \addrCollege of Computing, Georgia Institute of Technology 801 Atlantic Drive, Atlanta, GA 30332, USA
Abstract
In this paper, we consider an infinite dimensional exponential family $\mathcal{P}$ of probability densities, which are parametrized by functions in a reproducing kernel Hilbert space $\mathcal{H}$ , and show it to be quite rich in the sense that a broad class of densities on $\mathbb{R}^{d}$ can be approximated arbitrarily well in Kullback-Leibler (KL) divergence by elements in $\mathcal{P}$ . Motivated by this approximation property, the paper addresses the question of estimating an unknown density $p_{0}$ through an element in $\mathcal{P}$ . Standard techniques like maximum likelihood estimation (MLE) or pseudo MLE (based on the method of sieves), which are based on minimizing the KL divergence between $p_{0}$ and $\mathcal{P}$ , do not yield practically useful estimators because of their inability to efficiently handle the log-partition function. We propose an estimator $\hat{p}_{n}$ based on minimizing the Fisher divergence, $J(p_{0}\|p)$ between $p_{0}$ and $p\in\mathcal{P}$ , which involves solving a simple finite-dimensional linear system. When $p_{0}\in\mathcal{P}$ , we show that the proposed estimator is consistent, and provide a convergence rate of $n^{-\min\left\{\frac
中文速览
用再生核希尔伯特空间(RKHS)构造的无穷维指数族来拟合未知概率密度是个很自然的想法,但传统的最大似然估计因为要处理难以计算的对数配分函数(log-partition function)而几乎无法实用。本文转而最小化Fisher散度(Fisher divergence)——它不依赖配分函数——并证明这等价于求解一个简单的有限维线性方程组,从而得到一个计算上切实可行的密度估计量。理论上,当真实密度属于该模型族时,估计量是相合的,并给出了关于样本量 $n$ 的具体收敛速率;即使真实密度不在模型族内,估计量也能收敛到族内最优近似。数值实验表明,该方法在中高维情形下明显优于经典的核密度估计(KDE),且维度越高优势越突出,为高维非参数密度估计提供了一条兼顾理论严谨与工程可用的新路径。
原文 arXiv:1312.3516;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1312.3516v4