\scalerel*X MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
Xingyao Wang11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT, Zihan Wang1,2*†12absent†{}^{1,2*\dagger}start_FLOATSUPERSCRIPT 1 , 2 * † end_FLOATSUPERSCRIPT, Jiateng Liu11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT, Yangyi Chen11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT, Lifan Yuan1†1†{}^{1\dagger}start_FLOATSUPERSCRIPT 1 † end_FLOATSUPERSCRIPT, Hao Peng11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT, Heng Ji11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT 11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT University of Illinois Urbana-Champaign, 22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT Renmin University of China 11{}^{1}start_FLOATSUPERSCRIPT 1 Equal contribution. ††{}^{\dagger}start_FLOATSUPERSCRIPT † end_FLOATSUPERSCRIPTWork done during internship at UIUC.
Abstract
To solve complex tasks, large language models (LLMs) often require multiple rounds of interactions with the user, sometimes assisted by external tools. However, current evaluation protocols often emphasize benchmark performance with single-turn exchanges, neglecting the nuanced interactions among the user, LLMs, and external tools, while also underestimating the importance of natural language feedback from users. These oversights contribute to discrepancies between research benchmark evaluations and real-world use cases. We introduce MINT, a benchmark that evaluates LLMs’ ability to solve challenging tasks with multi-turn interactions by (1) using tools and (2) leveraging natural language feedback. To ensure reproducibility, we provide an evaluation framework where LLMs can access tools by executing Python code and receive users’ natural language feedback simulated by GPT-4. We repurpose a diverse set of established evaluation datasets focusing on reasoning, coding, and decision-making and carefully curate them into a compact subset for efficient evaluation. Our analysis of 20 open- and closed-source LLMs offers intriguing findings. (a) LLMs generally benefit from tools and languag
中文速览
大语言模型(LLM)在实际使用中往往需要与用户多轮交互、借助外部工具才能完成复杂任务,但现有评测基准几乎清一色只测"一问一答"的单轮表现,忽视了用户自然语言反馈和工具调用对模型能力的影响。为此,研究者提出了MINT这一多轮交互评测框架,让模型通过执行Python代码使用工具、并接收由GPT-4模拟的用户自然语言反馈,基于推理、代码生成和决策三类任务筛选出586个具有代表性的高难度样本来高效评估模型。对20个开源和闭源LLM的测试发现:所有模型都能从多轮工具使用(每轮提升1–8%)和自然语言反馈(提升2–17%)中获益;单轮表现更好并不保证多轮表现更好;最出人意料的是,监督指令微调(SIFT)和基于人类反馈的强化学习(RLHF)在大多数模型上反而损害了多轮交互能力。这一发现提示当前主流训练范式存在被忽视的短板,MINT有望成为推动LLM多轮交互能力研究的重要基准,尤其对缺乏大规模真实用户数据的开源社区意义显著。
原文 arXiv:2309.10691;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2309.10691v3