Ground-Truth Labels Matter: A Deeper Look into Input-Label Demonstrations
Kang Min Yoo∗#†‡§, Junyeob Kim∗§, Hyuhng Joon Kim§, Hyunsoo Cho§, Hwiyeol Jo‡, Sang-Woo Lee†‡♮, Sang-goo Lee§, Taeuk Kim#¶ §Seoul National University, †NAVER AI Lab, ‡NAVER CLOVA ♮Korea Advanced Institute of Science and Technology, ¶Hanyang University
Abstract
Despite recent explosion of interests in in-context learning, the underlying mechanism and the precise impact of the quality of demonstrations remain elusive. Intuitively, ground-truth labels should have as much impact in in-context learning (ICL) as supervised learning, but recent work reported that the input-label correspondence is significantly less important than previously thought. Intrigued by this counter-intuitive observation, we re-examine the importance of ground-truth labels in in-context learning. With the introduction of two novel metrics, namely Label-Correctness Sensitivity and Ground-truth Label Effect Ratio (GLER), we were able to conduct quantifiable analysis on the impact of ground-truth label demonstrations. Through extensive analyses, we find that the correct input-label mappings can have varying impacts on the downstream in-context learning performances, depending on the experimental configuration. Through additional studies, we identify key components, such as the verbosity of prompt templates and the language model size, as the controlling factor to achieve more noise-resilient ICL.
中文速览
大语言模型的"情境学习"(in-context learning, ICL)究竟有多依赖正确的示例标签?此前有研究声称标签正确与否对ICL性能影响甚微,但这篇论文对这一结论提出了质疑:作者发现在某些数据集和模型配置下,用完全错误标签替换正确标签会导致高达80%的精度下降,说明标签正确性的影响远比之前认为的更为显著。为了系统量化这一影响,作者提出了两个新指标——标签正确性敏感度(Label-Correctness Sensitivity)和真实标签效果比(GLER),并在17个分类数据集、多个语言模型(GPT-J、GPT-3)上进行了大量实验,发现标签正确性的影响因数据集、提示模板详略程度和模型规模的不同而差异显著。研究还进一步揭示出哪些因素(如更冗长的提示模板、更大的模型)能够增强ICL对噪声标签的鲁棒性,为在真实训练数据稀缺时如何更好地利用情境学习提供了实用指导。
原文 arXiv:2205.12685;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.12685v2