Superintelligence cannot be contained: Lessons from Computability Theory
Manuel Alfonseca Correspondence to: Escuela Politécnica Superior, Universidad Autónoma de Madrid, Madrid, Spain Manuel Cebrian Data61 Unit, Commonwealth Scientific and Industrial Research Organisation, Melbourne, Victoria, Australia Antonio Fernandez Anta IMDEA Networks Institute, Madrid, Spain Lorenzo Coviello Google, USA Andres Abeliuk Melbourne School of Engineering, University of Melbourne, Melbourne, Australia Iyad Rahwan The Media Lab, Massachusetts Institute of Technology, Cambridge, MA 02139, USA
Abstract
Superintelligence is a hypothetical agent that possesses intelligence far surpassing that of the brightest and most gifted human minds. In light of recent advances in machine intelligence, a number of scientists, philosophers and technologists have revived the discussion about the potential catastrophic risks entailed by such an entity. In this article, we trace the origins and development of the neo-fear of superintelligence, and some of the major proposals for its containment. We argue that such containment is, in principle, impossible, due to fundamental limits inherent to computing itself. Assuming that a superintelligence will contain a program that includes all the programs that can be executed by a universal Turing machine on input potentially as complex as the state of the world, strict containment requires simulations of such a program, something theoretically (and practically) infeasible.
中文速览
判断一个超级智能(superintelligence)是否会伤害人类,从根本上是无法用算法解决的——这篇文章正是要证明这一点。作者从图灵停机问题出发,将"超级智能遏制问题"转化为一个可计算性问题:如果存在一个程序能预判超级智能的所有有害行为,那它就能解决已被证明不可判定的停机问题,这是逻辑矛盾,因此这样的遏制程序根本不可能存在。换句话说,无论是把超级智能关进笼子、限制其能力,还是给它灌输人类友好的价值观,任何依赖程序化检测的控制策略在理论上都存在无法逾越的计算极限。这一结论对当前围绕人工智能存在风险的讨论具有重要意义:它表明我们不能指望通过某种"万能安全程序"来彻底遏制超级智能,相关的安全研究需要在这一根本性限制下重新审视其可行边界。
原文 arXiv:1607.00913;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1607.00913v1