Incomplete Contracting and AI Alignment
Dylan Hadfield-Menell Department of Electrical Engineering and Computer Science University of California Berkeley, CA 94720 OpenAI, Center for Human-Compatible Artificial Intelligence \AndGillian K. Hadfield Law School and Department of Economics University of Southern California Los Angeles, CA 90089 Center for Human-Compatible Artificial Intelligence
Abstract
There can never be complete communication between two people; a promise made and a promise heard are two different things…Thus [promises] can never be a complete basis for dealing with the future.
中文速览
当前的AI对齐问题(AI alignment)和经济学里研究了几十年的"不完全契约"问题本质上是同一类难题:无论是雇主给员工写合同,还是工程师给AI写奖励函数,都无法把所有情况和期望行为事先全部规定清楚。作者借助不完全契约理论的分析框架,系统梳理了AI奖励函数为何天然存在缺陷——包括设计者认知有限、无法预见所有情境、部分目标根本无法被精确编码等——并指出这种"错位"并非设计失误,而是不可避免的结构性问题。关键洞见在于:人类社会之所以能在合同不完整的情况下仍然顺畅合作,是因为有文化规范、法律默示条款、社会惯例等外部结构来自动填补合同漏洞;因此,作者提出AI对齐的研究方向应转向如何让AI学会调用类似的"外部结构"——即人类社会共享的价值背景——来弥补奖励函数的先天不足。这一视角将法律经济学的成熟工具引入AI安全领域,为超越"把奖励函数写得更完整"这一思路提供了全新的理论起点。
原文 arXiv:1804.04268;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1804.04268v1