Categorizing Variants of Goodhart’s Law
David Manheim Scott Garrabrant
Abstract
There are several distinct failure modes for overoptimization of systems on the basis of metrics. This occurs when a metric which can be used to improve a system is used to such an extent that further optimization is ineffective or harmful, and is sometimes termed Goodhart’s Law111As a historical note, Goodhart’s Law [1] as originally formulated states that “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.” This has been interpreted and explained more widely, perhaps to the point where it is ambiguous what the term means. Other closely related formulations, such as Campbell’s law (which arguably has scholarly precedence[3]) and the Lucas critique, were also initially specific, and their interpretation has also been expanded greatly. Lastly, the Cobra Effect and perverse incentives are often closely related to these failures, and the different effects interact. Because none of the terms were laid out formally, the categories proposed do not match what was originally discussed. A separate forthcoming paper intends to address the relationship between those formulations and the categories more formally explained here.. This
中文速览
用来指导系统优化的指标(proxy metric)与真实目标之间往往存在偏差,过度依赖指标优化会让系统越优化越偏离初衷——这一现象俗称"古德哈特定律"(Goodhart's Law),但学界对它的理解长期停留在模糊的定性层面。本文在 Garrabrant 早期框架的基础上,将这类失效模式细分为四大类并逐一建模:回归型(选指标时连噪声一起选进来)、极端型(优化把系统推进指标与目标关系失效的新区域)、因果型(监管者的干预行为本身切断了指标与目标的因果联系)、对抗型(存在与监管者目标不一致的博弈方主动操控指标)。研究结果表明,这四类失效机制在根源和应对方式上有实质差异,混为一谈会导致误判和错误应对。厘清这些机制对经济监管、公共政策制定乃至人工智能对齐(AI alignment)尤为关键——AI 带来的超强优化能力会成倍放大古德哈特效应的危害,使得精确区分失效类型比以往任何时候都更迫切。
原文 arXiv:1803.04585;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1803.04585v4