arXiv:2005.14165 · 中英对照阅读
语言模型是少样本学习者
Language Models are Few-Shot Learners
中文速览
传统语言模型做新任务通常要用成千上万条标注数据微调,难以像人一样靠几条示例或一句指令快速上手;GPT-3通过把自回归语言模型扩展到1750亿参数,并在推理时只用文字指令和上下文示例进行零样本、单样本或少样本学习,不更新模型参数。结果显示,模型规模越大,少样本能力整体越强,在翻译、问答、完形填空、算术、词语重组和临时理解新词等任务上表现突出,部分任务甚至接近或超过当时微调模型的最好成绩,还能生成让人难以分辨真假的新闻;但它在自然语言推理、部分阅读理解和数据污染等方面仍有明显问题。这项工作重要之处在于表明,单纯扩大模型规模可能让语言模型获得更通用的“看例子学任务”能力,推动了从任务专用微调走向更灵活的通用语言系统。
摘要
近期研究表明,在大规模文本语料库上进行预训练,随后针对特定任务进行微调,能够显著提升自然语言处理中的许多任务和基准测试的性能。尽管这种方法在架构上通常是任务无关的,但仍然需要包含数千乃至数万个示例的任务特定微调数据集。相比之下,人类通常仅凭少数示例或简单指令就能完成新的语言任务,而当前的自然语言处理系统在很大程度上仍难以做到这一点。本文表明,扩大语言模型的规模能够显著提升任务无关的少样本性能,有时甚至可以达到与此前最先进微调方法相当的竞争力。具体而言,我们训练了GPT-3,这是一种拥有1750亿参数的自回归语言模型,其参数量是此前任何非稀疏语言模型的10倍,并在少样本设置下测试其性能。对于所有任务,GPT-3都不进行任何梯度更新或微调,而是纯粹通过与模型的文本交互来指定任务和少样本示例。GPT-3在许多自然语言处理数据集上取得了出色的性能,包括翻译、问答和完形填空任务,以及若干要求即时推理或领域适应的任务,例如重新排列打乱的单词、在句子中使用新词,或进行三位数算术运算。与此同时,我们还发现,GPT-3的少样本学习在某些数据集上仍然表现不佳;此外,在一些数据集上,GPT-3还面临与在大规模网络语料库上训练相关的方法学问题。最后,我们发现,GPT-3能够生成新闻文章样本,使人类评估者难以将其与人类撰写的文章区分开来。我们讨论了这一发现以及GPT-3总体上可能带来的更广泛社会影响。
Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions – something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3’s few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general.
术语表
- NLP
- 自然语言处理
- pre-training
- 预训练
- fine-tuning
- 微调
- task-agnostic
- 任务无关
- few-shot learning
- 少样本学习
- zero-shot learning
- 零样本学习
- one-shot learning
- 单样本学习
- in-context learning
- 上下文学习
- meta-learning
- 元学习
- GPT-3
- GPT-3
- autoregressive language model
- 自回归语言模型
- transformer language model
- Transformer语言模型
- RNN
- 循环神经网络
- word vectors
- 词向量
- contextual state
- 上下文状态
- downstream task
- 下游任务
- task-specific architecture
- 任务特定架构
- task-specific dataset
- 任务特定数据集
- reading comprehension
- 阅读理解
- question answering
- 问答
- textual entailment
- 文本蕴含
- cloze task
- 完形填空任务
- Natural Questions
- Natural Questions
- CoQA
- CoQA
- log loss
- 对数损失
- out-of-distribution generalization
- 分布外泛化
- context window
- 上下文窗口
- few-shot demonstration
- 少样本示例
- domain adaptation
- 领域适应
- gradient update
- 梯度更新