Are Emergent Abilities in Large Language Models just In-Context Learning?
Sheng Lu1*, Irina Bigoulaeva1*, Rachneet Sachdeva1, Harish Tayyar Madabushi2 Iryna Gurevych1 1 Ubiquitous Knowledge Processing Lab, Technical University of Darmstadt 2 Department of Computer Science, The University of Bath www.ukp.tu-darmstadt.de
Abstract
Large language models, comprising billions of parameters and pre-trained on extensive web-scale corpora, have been claimed to acquire certain capabilities without having been specifically trained on them. These capabilities, referred to as “emergent abilities,” have been a driving force in discussions regarding the potentials and risks of language models. A key challenge in evaluating emergent abilities is that they are confounded by model competencies that arise through alternative prompting techniques, including in-context learning, which is the ability of models to complete a task based on a few examples. We present a novel theory that explains emergent abilities, taking into account their potential confounding factors, and rigorously substantiate this theory through over 1000 experiments. Our findings suggest that purported emergent abilities are not truly emergent, but result from a combination of in-context learning, model memory, and linguistic knowledge. Our work is a foundational step in explaining language model performance, providing a template for their efficient use and clarifying the paradox of their ability to excel in some instances while faltering in others. Thus,
中文速览
大型语言模型(LLM)被广泛认为能够在未经专门训练的情况下自发涌现出某些能力,尤其是推理、社交理解等功能性语言能力,这一"涌现能力"(emergent abilities)现象既引发了对模型潜力的乐观期待,也带来了对其潜在危险的担忧。研究者通过20个模型、22项任务、逾千次实验,系统检验了这些所谓涌现能力在排除上下文学习(in-context learning,ICL)干扰后是否依然存在。结果发现,当去掉ICL提示(即零样本条件下评估非指令微调模型)时,模型表现接近随机水平,根本谈不上真正的涌现;而那些看似涌现的能力,实际上源于ICL、模型记忆以及已有的形式语言知识三者的叠加效应。这一发现意义重大:它不仅澄清了"LLM在某些任务上出色、在另一些任务上失败"的看似矛盾现象,也说明当前阶段的LLM并不存在无法预测的潜在危险性功能能力,从而为更理性地评估和使用这类模型提供了理论依据。
原文 arXiv:2309.01809;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2309.01809v2