‘Indifference’ methods for managing agent rewards
Stuart Armstrong Future of Humanity Institute Oxford University UK、Xavier O’Rourke The Australian National University Canberra Australia Machine Intelligence Research Institute, Berkeley, USA.Future of Humanity Institute, Oxford University, UK.
Abstract
‘Indifference’ refers to a class of methods used to control reward based agents. Indifference techniques aim to achieve one or more of three distinct goals: rewards dependent on certain events (without the agent being motivated to manipulate the probability of those events), effective disbelief (where agents behave as if particular events could never happen), and seamless transition from one reward function to another (with the agent acting as if this change is unanticipated). This paper presents several methods for achieving these goals in the POMDP setting, establishing their uses, strengths, and requirements. These methods of control work even when the implications of the agent’s reward are otherwise not fully understood.
中文速览
在设计强化学习智能体时,有时我们希望它的奖励依赖于某些特定事件(比如顾客持有手环),但又不想让它因此产生操控这些事件的动机——例如文中那个会给所有人发手环以多卖酒的机器人。这篇论文系统梳理了"无差异"(indifference)这一类方法,将其核心目标归纳为三类:让奖励依赖某事件但不激励智能体操控该事件的概率、让智能体行为如同某事件永不会发生、以及让智能体在奖励函数切换时无缝过渡且事先不抱有预期。论文在部分可观测马尔可夫决策过程(POMDP)框架下,提出并比较了多种实现上述目标的具体构造方法,包括复合奖励、策略反事实、以及效用无差异等技术,并详细说明了各方法的适用场景、优势与前提条件。这项工作的价值在于,这些方法只需对奖励函数做相对简单的修改,即便设计者无法完全理解智能体行为的全部含义,也能对任意能力的智能体实施有效的安全控制,为AI对齐研究提供了可操作的技术工具。
原文 arXiv:1712.06365;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1712.06365v4