The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning
Xi Ye Greg Durrett Affiliation: Department of Computer Science Affiliation: The University of Texas at Austin Email:
Abstract
Does prompting a large language model (LLM) like GPT-3 with explanations improve in-context learning? We study this question on two NLP tasks that involve reasoning over text, namely question answering and natural language inference. We test the performance of four LLMs on three textual reasoning datasets using prompts that include explanations in multiple different styles. For these tasks, we find that including explanations in the prompts for OPT, GPT-3 (davinci), and InstructGPT (text-davinci-001) only yields small to moderate accuracy improvements over standard few-show learning. However, text-davinci-002 is able to benefit more substantially.
原文 arXiv:2205.03401;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2205.03401v2