Super-NaturalInstructions: A Benchmark of 1,600+ Language Tasks and Instructions
♢Yizhong Wang2 ♢Swaroop Mishra3 ♣Pegah Alipoormolabashi4 ♣Yeganeh Kordi5 Amirreza Mirzaei4 Anjana Arunkumar3 Arjun Ashok6 Arut Selvan Dhanasekaran3 Atharva Naik7 David Stap8 Eshaan Pathak9 Giannis Karamanolakis10 Haizhi Gary Lai11 Ishan Purohit12 Ishani Mondal13 Jacob Anderson3 Kirby Kuznia3 Krima Doshi3 Maitreya Patel3 Kuntal Kumar Pal3 Mehrad Moradshahi14 Mihir Parmar3 Mirali Purohit15 Neeraj Varshney3 Phani Rohitha Kaza3 Pulkit Verma3 Ravsehaj Singh Puri3 Rushang Karia3 Shailaja Keyur Sampat3 Savan Doshi3 Siddhartha Mishra16 Sujan Reddy17 Sumanta Patro18 Tanay Dixit19 Xudong Shen20 Chitta Baral3 Yejin Choi1,2 Noah A. Smith1,2 Hannaneh Hajishirzi1,2 Daniel Khashabi21 1Allen Institute for AI 2Univ. of Washington 3Arizona State Univ. 4Sharif Univ. of Tech. 5Tehran Polytechnic 6PSG College of Tech. 7IIT Kharagpur 8Univ. of Amsterdam 9UC Berkeley 10Columbia Univ. 11Factored AI 12Govt. Polytechnic Rajkot 13Microsoft Research 14Stanford Univ. 15Zycus Infotech 16Univ. of Massachusetts Amherst 17National Inst. of Tech. Karnataka 18TCS Research 19IIT Madras 20National Univ. of Singapore 21Johns Hopkins Univ.
Abstract
How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions,111Super-NaturalInstructions represents a super-sized expansion of NaturalInstructions Mishra et al. (2022b) which had 61 tasks. a benchmark of 1,616 diverse NLP tasks and their expert-written instructions. Our collection covers 76 distinct task types, including but not limited to classification, extraction, infilling, sequence tagging, text rewriting, and text composition. This large and diverse collection of tasks enables rigorous benchmarking of cross-task generalization under instructions—training models to follow instructions on a subset of tasks and evaluating them on the remaining unseen ones. Furthermore, we build T $k$ -Instruct, a transformer model trained to follow a variety of in-context instructions (plain language task definitions or $k$ -shot examples). Our experiments show that T $k$ -Instruct outperforms existing instruction-following models such as InstructGPT by over 9% on our benchmark despite being an order of magnitude smaller. We further analyze generalization as a function of various scal
中文速览
大规模、多样化的NLP指令数据集一直是推动模型泛化能力的关键瓶颈,为此研究者构建了包含1616个任务及配套专家撰写指令的Super-NaturalInstructions基准,覆盖76种任务类型和55种语言,规模远超以往同类数据集。基于此,他们训练了Tk-Instruct模型——通过多任务元训练让T5在已知任务的指令上学习"如何遵循指令",再迁移到从未见过的新任务上。实验表明,仅110亿参数的Tk-Instruct在119个未见英语任务上比1750亿参数的InstructGPT高出约10个ROUGE-L分,多语言版本的优势更达13分,人工评估也证实其77%的生成结果质量不亚于标准答案。这项工作的意义在于:它以完全公开的数据和相对小得多的模型复现并超越了封闭商业系统的跨任务泛化能力,为学术界提供了可复现、可扩展的研究基础。
原文 arXiv:2204.07705;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2204.07705v3