OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
Srinivasan Iyer Note: Equal contribution; alphabetical order. Xi Victoria Lin Ramakanth PasunuruTodor Mihaylov, Dániel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu,Punit Singh Koura, Xian Li, Brian O’Horo, Gabriel Pereyra , Jeff Wang,Christopher Dewan, Asli Celikyilmaz, Luke Zettlemoyer, Ves Stoyanov Meta AI Thanks: Work done while at Meta AI.
Abstract
Recent work has shown that fine-tuning large pre-trained language models on a collection of tasks described via instructions, a.k.a. instruction-tuning, improves their zero and few-shot generalization to unseen tasks. However, there is a limited understanding of the performance trade-offs of different decisions made during the instruction-tuning process. These decisions include the scale and diversity of the instruction-tuning benchmark, different task sampling strategies, fine-tuning with and without demonstrations, training using specialized datasets for reasoning and dialogue, and finally, the fine-tuning objectives themselves. In this paper, we characterize the effect of instruction-tuning decisions on downstream task performance when scaling both model and benchmark sizes. To this end, we create OPT-IML Bench: a large benchmark for Instruction Meta-Learning (IML) of 2000 NLP tasks consolidated into task categories from 8 existing benchmarks, and prepare an evaluation framework to measure three types of model generalizations: to tasks from fully held-out categories, to held-out tasks from seen categories, and to held-out instances from seen tasks. Through the lens of this frame
原文 arXiv:2212.12017;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2212.12017v3