How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
Yizhong Wang ♣♠ Hamish Ivison††♣ Pradeep Dasigi♣ Jack Hessel♣ Tushar Khot♣ Khyathi Raghavi Chandu♣ David Wadden♣ Kelsey MacMillan♣ Noah A. Smith♣♠ Iz Beltagy♣ Hannaneh Hajishirzi♣♠ ♣Allen Institute for AI ♠University of Washington Equal contribution.
Abstract
In this work we explore recent advances in instruction-tuning language models on a range of open instruction-following datasets. Despite recent claims that open models can be on par with state-of-the-art proprietary models, these claims are often accompanied by limited evaluation, making it difficult to compare models across the board and determine the utility of various resources. We provide a large set of instruction-tuned models from 6.7B to 65B parameters in size, trained on 12 instruction datasets ranging from manually curated (e.g., OpenAssistant) to synthetic and distilled (e.g., Alpaca) and systematically evaluate them on their factual knowledge, reasoning, multilinguality, coding, safety, and open-ended instruction following abilities through a collection of automatic, model-based, and human-based metrics. We further introduce Tülu , our best performing instruction-tuned model suite finetuned on a combination of high-quality open resources.
原文 arXiv:2306.04751;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2306.04751v2