A Formal Solution to the Grain of Truth Problem
Jan Leike Australian National University、Jessica Taylor Machine Intelligence Research Inst.、Benya Fallenstein Machine Intelligence Research Inst.
Abstract
A Bayesian agent acting in a multi-agent environment learns to predict the other agents’ policies if its prior assigns positive probability to them (in other words, its prior contains a grain of truth). Finding a reasonably large class of policies that contains the Bayes-optimal policies with respect to this class is known as the grain of truth problem. Only small classes are known to have a grain of truth and the literature contains several related impossibility results. In this paper we present a formal and general solution to the full grain of truth problem: we construct a class of policies that contains all computable policies as well as Bayes-optimal policies for every lower semicomputable prior over the class. When the environment is unknown, Bayes-optimal agents may fail to act optimally even asymptotically. However, agents based on Thompson sampling converge to play $\varepsilon$ -Nash equilibria in arbitrary unknown computable multi-agent environments. While these results are purely theoretical, we show that they can be computationally approximated arbitrarily closely.
中文速览
多智能体强化学习中,若每个智能体的先验对其他智能体的策略赋予正概率(即"真实一粒"条件),贝叶斯最优智能体就能学会预测对手并收敛到纳什均衡——但如何找到一个足够大、同时又对自身贝叶斯最优策略封闭的策略类,长期以来是公开难题。本文借助"反射性神谕"(reflective oracle)这一可自我引用的概率神谕,构造出一个包含所有可计算策略、且关于该类上任意下半可计算先验的贝叶斯最优策略也都在其中的策略类,从而在一般意义上完整解决了这一问题。进一步地,当环境未知时,纯贝叶斯最优智能体可能因探索不足而永远无法学到最优行为,但采用汤普森采样(Thompson sampling)的智能体凭借其内在随机性能充分探索,最终在任意未知可计算多智能体环境中收敛到 ε-纳什均衡。该框架虽为理论结果,但所依赖的反射性神谕被证明是极限可计算的,因而整套方法可以被任意精度地计算近似,具有重要的理论奠基意义。
原文 arXiv:1609.05058;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1609.05058v1