Predictability and Surprise in Large Generative Models
Deep Ganguli , Danny Hernandez , Liane Lovitt , Nova DasSarma , Tom Henighan , Andy Jones , Nicholas Joseph , Jackson Kernion , Ben Mann , Amanda Askell , Yuntao Bai , Anna Chen , Tom Conerly , Dawn Drain , Nelson Elhage , Sheer El Showk , Stanislav Fort , Zac Hatfield-Dodds , Scott Johnston , Shauna Kravec , Neel Nanda , Kamal Ndousse , Catherine Olsson , Daniela Amodei , Tom Brown , Jared Kaplan , Sam McCandlish , Chris Olah , Dario Amodei and Jack Clark AnthropicSan FranciscoUSA
Abstract
Large-scale pre-training has recently emerged as a technique for creating capable, general-purpose, generative models such as GPT-3, Megatron-Turing NLG, Gopher, and many others. In this paper, we highlight a counterintuitive property of such models and discuss the policy implications of this property. Namely, these generative models have a paradoxical combination of predictable loss on a broad training distribution (as embodied in their "scaling laws"), and unpredictable specific capabilities, inputs, and outputs. We believe that the high-level predictability and appearance of useful capabilities drives rapid development of such models, while the unpredictable qualities make it difficult to anticipate the consequences of model deployment. We go through examples of how this combination can lead to socially harmful behavior with examples from the literature and real world observations, and we also perform two novel experiments to illustrate our point about harms from unpredictability. Furthermore, we analyze how these conflicting properties combine to give model developers various motivations for deploying these models, and challenges that can hinder deployment. We conclude with a l
中文速览
大规模预训练语言模型(如GPT-3)遵循"规模定律"(scaling laws),即投入更多算力和数据就能可预测地提升模型整体表现,这种可预测性吸引了大量机构竞相投入开发;然而与此同时,模型在特定任务上的能力何时涌现、会产生哪些具体输出,却几乎无法提前预判。作者将这一"宏观可预测、微观不可预测"的矛盾特性定义为大型生成模型的核心悖论,并通过文献梳理和两项原创实验展示了这种不可预测性如何导致有害内容随规模扩大而涌现。基于此,论文分析了经济利益、学术声誉等多重驱动力为何让开发者在风险未明的情况下仍持续推进模型部署,并最终向政策制定者、研究者和资助方提出了一系列干预建议。这项研究的价值在于,它将规模定律与社会风险这两个通常分开讨论的议题整合进同一个分析框架,为理解和规范新一代AI系统提供了更完整的视角。
原文 arXiv:2202.07785;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2202.07785v2