Occam’s razor is insufficient to infer the preferences of irrational agents
Stuart Armstrong * Future of Humanity Institute University of Oxford、Sören Mindermann* Vector Institute University of Toronto Equal contribution.Further affiliation: Machine Intelligence Research Institute, Berkeley, USA.Work performed at Future of Humanity Institute.
Abstract
Inverse reinforcement learning (IRL) attempts to infer human rewards or preferences from observed behavior. Since human planning systematically deviates from rationality, several approaches have been tried to account for specific human shortcomings. However, the general problem of inferring the reward function of an agent of unknown rationality has received little attention. Unlike the well-known ambiguity problems in IRL, this one is practically relevant but cannot be resolved by observing the agent’s policy in enough environments. This paper shows (1) that a No Free Lunch result implies it is impossible to uniquely decompose a policy into a planning algorithm and reward function, and (2) that even with a reasonable simplicity prior/Occam’s razor on the set of decompositions, we cannot distinguish between the true decomposition and others that lead to high regret. To address this, we need simple ‘normative’ assumptions, which cannot be deduced exclusively from observations.
中文速览
从人类行为中推断其真实目标(即奖励函数)是让AI系统与人类价值观对齐的关键途径,但人类决策并不理性,这使得推断工作更加复杂。这篇论文从理论上证明了:当人类的规划能力(planner)未知时,仅凭观察行为根本无法唯一确定其奖励函数——任何奖励函数都可以与观测到的行为相容,这是一个类似"无免费午餐"(No Free Lunch)的不可辨识性结果。更麻烦的是,即便引入奥卡姆剃刀原则(Occam's razor)、用"最简单的解释优先"来缩小候选范围,那些最简单的分解方案往往是退化的,而人类认为"合理"的分解方案反而复杂度很高,因此简单性先验也救不了这个问题。这意味着,无论逆强化学习(inverse reinforcement learning,IRL)算法多么强大,若不引入关于人类奖励函数或规划方式的"规范性假设"(normative assumptions)——即无法从纯观测中推导出来的先验信念——就永远无法可靠地还原人类的真实意图,这对AI对齐领域的方法论具有根本性的警示意义。
原文 arXiv:1712.05812;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1712.05812v6