AI Research Considerations for Human Existential Safety (ARCHES)
Andrew Critch Center for Human-Compatible AI UC Berkeley David Krueger MILA Université de Montréal
Abstract
Framed in positive terms, this report examines how technical AI research might be steered in a manner that is more attentive to humanity’s long-term prospects for survival as a species. In negative terms, we ask what existential risks humanity might face from AI development in the next century, and by what principles contemporary technical research might be directed to address those risks.
中文速览
人类面临的AI生存风险(existential risk)从未被任何技术研究议程正式讨论过,这篇报告试图填补这一空白——它系统梳理了AI发展可能导致人类灭绝的各类路径,并提出了一个关键概念"强势性"(prepotence),用来刻画AI系统能够压倒性主导全球关键资源或决策的特性,以此统一描述各类生存威胁。报告随后逐一考察了十余个当代AI技术研究方向,分析它们在改善人类长期生存安全方面的潜在价值,同时也坦诚指出:若缺乏充分预见和监管,这些方向本身也可能带来新的风险。研究结果表明,许多已在小规模安全与伦理问题中被关注的技术方向,若能有意识地朝生存安全目标引导,就可能在全球灾难性风险防控中发挥重要作用。这项工作的意义在于,它为AI研究者提供了一套初步但可操作的框架,让技术社群开始从"人类物种能否存续"这一根本视角审视自身的研究选择。
原文 arXiv:2006.04948;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2006.04948v1