Overlap matrix concentration in optimal Bayesian inference
Jean Barbier
Abstract
We consider models of Bayesian inference of signals with vectorial components of finite dimensionality. We show that under a proper perturbation these models are replica symmetric in the sense that the overlap matrix concentrates. The overlap matrix is the order parameter in these models and is directly related to error metrics such as minimum mean-square errors. Our proof is valid in the optimal Bayesian inference setting. This means that it relies on the assumption that the model and all its hyper-parameters are known so that the posterior distribution can be written exactly. Examples of important problems in high-dimensional inference and learning to which our results apply are low-rank tensor factorization, the committee machine neural network with a finite number of hidden neurons in the teacher-student scenario, or multi-layer versions of the generalized linear model.
中文速览
在高维贝叶斯推断(Bayesian inference)中,刻画推断质量的核心量叫做"重叠矩阵"(overlap matrix),它直接决定了最小均方误差等关键性能指标,但当信号的每个分量本身是一个向量时,这个重叠量从标量升级为矩阵,传统的集中性证明方法不再适用。本文针对这类向量信号的最优贝叶斯推断模型,通过引入精心设计的微小扰动并结合"西森利恒等式"(Nishimori identities)——一组由贝叶斯公式天然保证的统计恒等式——严格证明了重叠矩阵在整个参数空间上均集中于确定性值,即模型满足"复制对称性"(replica symmetry)。这一结论适用于一大类重要模型,包括低秩张量分解、有限隐神经元的委员会机(committee machine)神经网络的师生场景,以及多层广义线性模型(multi-layer generalized linear model)。该结果填补了物理学启发的replica方法与严格数学证明之间长期存在的空白,为上述模型自由能和互信息的渐近公式提供了坚实的理论基础。
原文 arXiv:1904.02808;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1904.02808v2