Rethinking the Effective Sample Size
Víctor Elvira University of Edinburgh, United Kingdom The Alan Turing Institute, United Kingdom Luca Martino King Juan Carlos University of Madrid, Spain Christian P. Robert Université Paris Dauphine, France
Abstract
The effective sample size (ESS) is widely used in sample-based simulation methods for assessing the quality of a Monte Carlo approximation of a given distribution and of related integrals. In this paper, we revisit the approximation of the ESS in the specific context of importance sampling (IS). The derivation of this approximation, that we will denote as $\widehat{\text{ESS}}$ , is partially available in [27]. This approximation has been widely used in the last 25 years due to its simplicity as a practical rule of thumb in a wide variety of importance sampling methods. However, we show that the multiple assumptions and approximations in the derivation of $\widehat{\text{ESS}}$ , makes it difficult to be considered even as a reasonable approximation of the ESS. We extend the discussion of the $\widehat{\text{ESS}}$ in the multiple importance sampling (MIS) setting, we display numerical examples, and we discuss several avenues for developing alternative metrics. This paper does not cover the use of ESS for MCMC algorithms.
中文速览
重新审视在重要性采样(importance sampling, IS)中被广泛使用了近三十年的"有效样本量"近似公式($\widehat{\text{ESS}} = 1/\sum \bar{w}_n^2$),发现这个公式在推导过程中叠加了多个强假设和粗糙近似,包括忽略偏差、两次使用delta方法展开、以及一个与被积函数$h$完全无关的化简,导致最终结果与真实ESS之间存在难以控制的系统性偏差。论文逐步还原了从1992年Kong等人技术报告到该公式的完整推导链条,明确指出每一步近似的条件和代价,并通过数值实验展示了该公式在极端情形下给出严重误导性结论的场景,例如当IS估计器方差实际上小于直接蒙特卡洛估计器时,公式却始终给出小于$N$的值。这项工作的意义在于为大量依赖$\widehat{\text{ESS}}$进行算法诊断和比较的研究者敲响警钟,并指出开发更可靠替代度量的若干方向,对粒子滤波、自适应IS等众多方法的实践具有直接参考价值。
原文 arXiv:1809.04129;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1809.04129v2