GTA: A Benchmark for General Tool Agents
Jize Wang Zerun Ma Yining Li Songyang Zhang Affiliation: Shanghai Jiao Tong University Shanghai AI Affiliation: Shanghai Jiao Tong University Shanghai AI Affiliation: Shanghai Jiao Tong University Shanghai AI Affiliation: Shanghai Jiao Tong University Shanghai AI Cailian Chen Kai Chen Xinyi Le Thanks: Corresponding Authors.
Abstract
Significant focus has been placed on integrating large language models (LLMs) with various tools in developing general-purpose agents. This poses a challenge to LLMs’ tool-use capabilities. However, there are evident gaps between existing tool-use evaluations and real-world scenarios. Current evaluations often use AI-generated queries, single-step tasks, dummy tools, and text-only interactions, failing to effectively reveal the agents’ real-world problem-solving abilities. To address this, we propose GTA, a benchmark for General Tool Agents, featuring three main aspects: (i) Real user queries: human-written queries with simple real-world objectives but implicit tool-use, requiring the LLM to reason the suitable tools and plan the solution steps. (ii) Real deployed tools: an evaluation platform equipped with tools across perception, operation, logic, and creativity categories to evaluate the agents’ actual task execution performance. (iii) Real multimodal inputs: authentic image files, such as spatial scenes, web page screenshots, tables, code snippets, and printed/handwritten materials, used as the query contexts to align with real-world scenarios closely. We design 229 real-world
原文 arXiv:2407.08713;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2407.08713v2