Conservative Agency
Alexander Matt Turner1 Dylan Hadfield-Menell2、Prasad Tadepalli1 1 Oregon State University 2 UC Berkeley {turneale,
Abstract
Reward functions are easy to misspecify; although designers can make corrections after observing mistakes, an agent pursuing a misspecified reward function can irreversibly change the state of its environment. If that change precludes optimization of the correctly specified reward function, then correction is futile. For example, a robotic factory assistant could break expensive equipment due to a reward misspecification; even if the designers immediately correct the reward function, the damage is done. To mitigate this risk, we introduce an approach that balances optimization of the primary reward function with preservation of the ability to optimize auxiliary reward functions. Surprisingly, even when the auxiliary reward functions are randomly generated and therefore uninformative about the correctly specified reward function, this approach induces conservative, effective behavior.
中文速览
奖励函数(reward function)写错了怎么办——这是强化学习走向现实世界时最棘手的问题之一:智能体在设计者发现并纠正错误之前,可能已经对环境造成了不可挽回的破坏。为此,研究者提出了"可达效用保留"(Attainable Utility Preservation,AUP)方法:在让智能体优化主要奖励函数的同时,额外惩罚它对一组辅助奖励函数的可优化程度所造成的改变,从而让智能体倾向于采取保守、可逆的行动。最令人惊讶的发现是,即使这些辅助奖励函数完全随机生成、与真实目标毫无关联,AUP智能体依然能在完成主任务的同时自发避免副作用、减少对环境的不可逆改动。这一结果表明,无需事先知道"哪些事情不该做",只要保住对任意目标的优化能力,就能大概率保住对正确目标的优化能力,为解决奖励误设计问题提供了一条简洁而通用的路径。
原文 arXiv:1902.09725;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1902.09725v3