Aha.
正在载入中英对照阅读…

arXiv:2005.14165 · 中英对照阅读

语言模型是少样本学习者

Language Models are Few-Shot Learners

Tom B. Brown、Benjamin Mann、Nick Ryder、Melanie Subbiah、Jared Kaplan、Prafulla Dhariwal、Arvind Neelakantan、Pranav Shyam、Girish Sastry、Amanda Askell、Sandhini Agarwal、Ariel Herbert-Voss、Gretchen Krueger、Tom Henighan、Rewon Child、Aditya Ramesh、Daniel M. Ziegler、Jeffrey Wu、Clemens Winter、Christopher Hesse、Mark Chen、Eric Sigler、Mateusz Litwin、Scott Gray、Benjamin Chess、Jack Clark、Christopher Berner、Sam McCandlish、Alec Radford、Ilya Sutskever、Dario Amodei、OpenAI

中文速览

传统语言模型做新任务通常要用成千上万条标注数据微调,难以像人一样靠几条示例或一句指令快速上手;GPT-3通过把自回归语言模型扩展到1750亿参数,并在推理时只用文字指令和上下文示例进行零样本、单样本或少样本学习,不更新模型参数。结果显示,模型规模越大,少样本能力整体越强,在翻译、问答、完形填空、算术、词语重组和临时理解新词等任务上表现突出,部分任务甚至接近或超过当时微调模型的最好成绩,还能生成让人难以分辨真假的新闻;但它在自然语言推理、部分阅读理解和数据污染等方面仍有明显问题。这项工作重要之处在于表明,单纯扩大模型规模可能让语言模型获得更通用的“看例子学任务”能力,推动了从任务专用微调走向更灵活的通用语言系统。

摘要

近期研究表明,在大规模文本语料库上进行预训练,随后针对特定任务进行微调,能够显著提升自然语言处理中的许多任务和基准测试的性能。尽管这种方法在架构上通常是任务无关的,但仍然需要包含数千乃至数万个示例的任务特定微调数据集。相比之下,人类通常仅凭少数示例或简单指令就能完成新的语言任务,而当前的自然语言处理系统在很大程度上仍难以做到这一点。本文表明,扩大语言模型的规模能够显著提升任务无关的少样本性能,有时甚至可以达到与此前最先进微调方法相当的竞争力。具体而言,我们训练了GPT-3,这是一种拥有1750亿参数的自回归语言模型,其参数量是此前任何非稀疏语言模型的10倍,并在少样本设置下测试其性能。对于所有任务,GPT-3都不进行任何梯度更新或微调,而是纯粹通过与模型的文本交互来指定任务和少样本示例。GPT-3在许多自然语言处理数据集上取得了出色的性能,包括翻译、问答和完形填空任务,以及若干要求即时推理或领域适应的任务,例如重新排列打乱的单词、在句子中使用新词,或进行三位数算术运算。与此同时,我们还发现,GPT-3的少样本学习在某些数据集上仍然表现不佳;此外,在一些数据集上,GPT-3还面临与在大规模网络语料库上训练相关的方法学问题。最后,我们发现,GPT-3能够生成新闻文章样本,使人类评估者难以将其与人类撰写的文章区分开来。我们讨论了这一发现以及GPT-3总体上可能带来的更广泛社会影响。

Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions – something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3’s few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general.

术语表

NLP
自然语言处理
pre-training
预训练
fine-tuning
微调
task-agnostic
任务无关
few-shot learning
少样本学习
zero-shot learning
零样本学习
one-shot learning
单样本学习
in-context learning
上下文学习
meta-learning
元学习
GPT-3
GPT-3
autoregressive language model
自回归语言模型
transformer language model
Transformer语言模型
RNN
循环神经网络
word vectors
词向量
contextual state
上下文状态
downstream task
下游任务
task-specific architecture
任务特定架构
task-specific dataset
任务特定数据集
reading comprehension
阅读理解
question answering
问答
textual entailment
文本蕴含
cloze task
完形填空任务
Natural Questions
Natural Questions
CoQA
CoQA
log loss
对数损失
out-of-distribution generalization
分布外泛化
context window
上下文窗口
few-shot demonstration
少样本示例
domain adaptation
领域适应
gradient update
梯度更新