ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
Haiyang Shen1, Yue Li3, Desong Meng4, Dongqi Cai5, Sheng Qi2, Li Zhang5, Mengwei Xu5, Yun Ma1 1Institute for Artificial Intelligence, Peking University 2School of Computer Science, Peking University 3School of Software、Microelectronics, Peking University 4School of Electronics Engineering and Computer Science, Peking University 5Beijing University of Posts and Telecommunications Corresponding author
Abstract
Recent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recent work demonstrates that these API-based agents exhibit relatively strong autonomy and planning capabilities. However, their ability to handle multi-dimensional difficulty levels, diverse task types, and real-world demands remains unknown. In this paper, we introduce ShortcutsBench, a large-scale benchmark for the comprehensive evaluation of API-based agents in solving real-world complex tasks. ShortcutsBench includes a wealth of real APIs from Apple Inc., refined user queries, human-annotated high-quality action sequences, detailed parameter filling values, and parameters requesting necessary input from the system or user. We revealed how existing benchmarks / datasets struggle to accommodate the advanced reasoning capabilities of existing more intelligent LLMs. Moreover, our extensive evaluation of agents built with $5$ leading open-source (size $\geq$ 57B) and $5$ closed-source LLMs (e.g. Gemini-1.5-Pro and GPT-4o-mini) with varying intelligence level reveals significant limitations of existing API-bas
中文速览
现有评测基准普遍存在API数量少、任务过于简单的问题,导致连参数量仅30亿的小模型都能轻松达到高分,根本无法区分不同大语言模型(LLM)的真实能力。为此,研究团队从苹果"快捷指令"(Shortcuts)平台抓取真实工作流数据,构建了ShortcutsBench——一个涵盖88款应用、1414个真实API、7627条用户快捷指令的大规模评测基准,同时覆盖API选择、参数填充、以及向系统或用户索取缺失信息这三个完整环节。他们用5个开源大模型(规模≥57B)和5个闭源大模型(如GPT-4o-mini、Gemini-1.5-Pro)进行全面测评,发现所有模型在复杂任务上仍存在明显短板:从查询中提取必要参数最为困难,而主动向用户索取关键信息的意识则普遍严重不足。这项工作揭示了当前基于API的智能体在真实复杂场景下的能力天花板,为推动LLM智能体走向实际部署提供了重要参照。
原文 arXiv:2407.00132;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2407.00132v3