Is Power-Seeking AI an Existential Risk?
Joseph Carlsmith Open Philanthropy April 2021 Shorter version | Video presentation | Slides | Audio version | Reviews (including superforecasters)
Abstract
This report examines what I see as the core argument for concern about existential risk from misaligned artificial intelligence. I proceed in two stages. First, I lay out a backdrop picture that informs such concern. On this picture, intelligent agency is an extremely powerful force, and creating agents much more intelligent than us is playing with fire – especially given that if their objectives are problematic, such agents would plausibly have instrumental incentives to seek power over humans. Second, I formulate and evaluate a more specific six-premise argument that creating agents of this kind will lead to existential catastrophe by 2070. On this argument, by 2070: (1) it will become possible and financially feasible to build relevantly powerful and agentic AI systems; (2) there will be strong incentives to do so; (3) it will be much harder to build aligned (and relevantly powerful/agentic) AI systems than to build misaligned (and relevantly powerful/agentic) AI systems that are still superficially attractive to deploy; (4) some such misaligned systems will seek power over humans in high-impact ways; (5) this problem will scale to the full disempowerment of humanity; and (6) su
中文速览
有研究者担心超级智能AI会夺权并毁掉人类的长远未来,但这种担忧到底有多充分、逻辑链条究竟在哪里断裂或成立,长期缺乏系统性的论证与概率量化。作者将这一担忧拆解为六个可独立评估的前提:到2070年,强大的"APS系统"(高级能力、规划性、战略意识兼备的AI)在技术和经济上将成为可能;会有强烈动机去构建它;构建"目标对齐"的版本将远难于构建"表面看起来可用但目标有问题"的版本;部分存在目标偏差的系统会以高影响方式争夺对人类的控制权;这种权力争夺会累积到永久剥夺全人类自主权的程度;而这种剥夺本身就构成存在性灾难。作者逐条分析每个前提的可信度并赋予主观概率,最终得出到2070年此类存在性灾难发生概率约为5%(2022年修订后升至>10%)的估计。这项工作的价值不在于给出一个精确数字,而在于把一个长期笼罩在直觉和隐喻中的风险论证透明化,使其每个环节都能被独立审视和反驳,从而为严肃的学术讨论提供可操作的框架。
原文 arXiv:2206.13353;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2206.13353v2