On the Measure of Intelligence
François Chollet Google, Inc. I thank José Hernández-Orallo, Julian Togelius, Christian Szegedy, and Martin Wicke for their valuable comments on the draft of this document.
Abstract
To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons between two systems, as well as comparisons with humans. Over the past hundred years, there has been an abundance of attempts to define and measure intelligence, across both the fields of psychology and AI. We summarize and critically assess these definitions and evaluation approaches, while making apparent the two historical conceptions of intelligence that have implicitly guided them. We note that in practice, the contemporary AI community still gravitates towards benchmarking intelligence by comparing the skill exhibited by AIs and humans at specific tasks, such as board games and video games. We argue that solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience: unlimited priors or unlimited training data allow experimenters to “buy” arbitrary levels of skills for a system, in a way that masks the system’s own generalization power. We then articulate a
中文速览
人工智能领域长期缺乏一个精确、可操作的"智能"定义,导致研究者只能靠让AI在象棋、视频游戏等特定任务上击败人类来衡量进展,但这种做法实际上度量的是"技能"而非真正的智能——因为只要堆够训练数据和先验知识,任何系统都能在某项任务上表现出色,这掩盖了系统真正的泛化能力。作者基于算法信息论提出了一个新的正式定义:智能是"技能习得效率"(skill-acquisition efficiency),即在有限先验知识和有限经验的条件下,系统能多高效地掌握新技能。依据这一定义,作者制定了通用智能基准应满足的设计准则,并据此构建了一个名为"抽象与推理语料库"(Abstraction and Reasoning Corpus,ARC)的新基准测试,其核心先验知识被刻意设计得与人类先天认知能力尽量吻合。ARC的意义在于,它让我们第一次拥有了一把能公平比较AI与人类"流体智力"(fluid intelligence)的尺子,有望把AI研究从刷分竞赛真正引向对通用智能的追求。
原文 arXiv:1911.01547;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1911.01547v2