Super-NaturalInstructions: A Benchmark of 1,600+ Language Tasks and Instructions
♢Yizhong Wang2 ♢Swaroop Mishra3 ♣Pegah Alipoormolabashi4 ♣Yeganeh Kordi5 Amirreza Mirzaei4 Anjana Arunkumar3 Arjun Ashok6 Arut Selvan Dhanasekaran3 Atharva Naik7 David Stap8 Eshaan Pathak9 Giannis Karamanolakis10 Haizhi Gary Lai11 Ishan Purohit12 Ishani Mondal13 Jacob Anderson3 Kirby Kuznia3 Krima Doshi3 Maitreya Patel3 Kuntal Kumar Pal3 Mehrad Moradshahi14 Mihir Parmar3 Mirali Purohit15 Neeraj Varshney3 Phani Rohitha Kaza3 Pulkit Verma3 Ravsehaj Singh Puri3 Rushang Karia3 Shailaja Keyur Sampat3 Savan Doshi3 Siddhartha Mishra16 Sujan Reddy17 Sumanta Patro18 Tanay Dixit19 Xudong Shen20 Chitta Baral3 Yejin Choi1,2 Noah A. Smith1,2 Hannaneh Hajishirzi1,2 Daniel Khashabi21 1Allen Institute for AI 2Univ. of Washington 3Arizona State Univ. 4Sharif Univ. of Tech. 5Tehran Polytechnic 6PSG College of Tech. 7IIT Kharagpur 8Univ. of Amsterdam 9UC Berkeley 10Columbia Univ. 11Factored AI 12Govt. Polytechnic Rajkot 13Microsoft Research 14Stanford Univ. 15Zycus Infotech 16Univ. of Massachusetts Amherst 17National Inst. of Tech. Karnataka 18TCS Research 19IIT Madras 20National Univ. of Singapore 21Johns Hopkins Univ.
Abstract
How well can NLP models generalize to a variety of unseen tasks when provided with task instructions? To address this question, we first introduce Super-NaturalInstructions,111Super-NaturalInstructions represents a super-sized expansion of NaturalInstructions Mishra et al. (2022b) which had 61 tasks. a benchmark of 1,616 diverse NLP tasks and their expert-written instructions. Our collection covers 76 distinct task types, including but not limited to classification, extraction, infilling, sequence tagging, text rewriting, and text composition. This large and diverse collection of tasks enables rigorous benchmarking of cross-task generalization under instructions—training models to follow instructions on a subset of tasks and evaluating them on the remaining unseen ones. Furthermore, we build T $k$ -Instruct, a transformer model trained to follow a variety of in-context instructions (plain language task definitions or $k$ -shot examples). Our experiments show that T $k$ -Instruct outperforms existing instruction-following models such as InstructGPT by over 9% on our benchmark despite being an order of magnitude smaller. We further analyze generalization as a function of various scal
原文 arXiv:2204.07705;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2204.07705v3