ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
Haiyang Shen Affiliation: Institute for Artificial Intelligence, Peking University Yue Li Affiliation: School of Software、Microelectronics, Peking University Desong Meng Affiliation: School of Electronics Engineering and Computer Science, Peking University Dongqi Cai Affiliation: Beijing University of Posts and Sheng Qi Affiliation: School of Computer Science, Peking University Li Zhang Affiliation: Beijing University of Posts and Mengwei Xu Affiliation: Beijing University of Posts and Yun Ma Thanks: Corresponding author Affiliation: Institute for Artificial Intelligence, Peking University
Abstract
Recent advancements in integrating large language models (LLMs) with application programming interfaces (APIs) have gained significant interest in both academia and industry. Recent work demonstrates that these API-based agents exhibit relatively strong autonomy and planning capabilities. However, their ability to handle multi-dimensional difficulty levels, diverse task types, and real-world demands remains unknown. In this paper, we introduce ShortcutsBench, a large-scale benchmark for the comprehensive evaluation of API-based agents in solving real-world complex tasks. ShortcutsBench includes a wealth of real APIs from Apple Inc., refined user queries, human-annotated high-quality action sequences, detailed parameter filling values, and parameters requesting necessary input from the system or user. We revealed how existing benchmarks / datasets struggle to accommodate the advanced reasoning capabilities of existing more intelligent LLMs. Moreover, our extensive evaluation of agents built with $5$ leading open-source (size $\geq$ 57B) and $5$ closed-source LLMs (e.g. Gemini-1.5-Pro and GPT-4o-mini) with varying intelligence level reveals significant limitations of existing API-bas
原文 arXiv:2407.00132;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2407.00132v3