Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub
Bohan Lyu11 1 Equal contribution. Affiliation: Department of Computer Science and Technology, Tsinghua University Xin Cong11 1 Equal contribution.22 2 Corresponding author. Affiliation: Department of Computer Science and Technology, Tsinghua University Heyang Yu Affiliation: Department of Computer Science and Technology, Tsinghua University Pan Yang Affiliation: Department of Computer Science and Technology, Tsinghua University Cheng Qian Affiliation: Department of Computer Science and Technology, Tsinghua University Affiliation: University of Illinois Urbana-Champaign Zihe Wang Affiliation: Department of Computer Science and Technology, Tsinghua University Yujia Qin Affiliation: Department of Computer Science and Technology, Tsinghua University Yining Ye Affiliation: Department of Computer Science and Technology, Tsinghua University Yaxi Lu Affiliation: Department of Computer Science and Technology, Tsinghua University Chen Qian Affiliation: Department of Computer Science and Technology, Tsinghua University Affiliation: School of Artificial Intelligence, Shanghai Jiao Tong Zhong Zhang Affiliation: Department of Computer Science and Technology, Tsinghua University Yukun Yan Affiliation: Department of Computer Science and Technology, Tsinghua University Yankai Lin Affiliation: Gaoling School of Artificial Intelligence, Renmin University of China Zhiyuan Liu22 2 Corresponding author. Affiliation: Department of Computer Science and Technology, Tsinghua University Maosong Sun Affiliation: Department of Computer Science and Technology, Tsinghua University
Abstract
Large Language Models (LLMs) excel in traditional natural language processing tasks but struggle with problems that require complex domain-specific calculations or simulations. While equipping LLMs with external tools to build LLM-based agents can enhance their capabilities, existing approaches lack the flexibility to address diverse and ever-evolving user queries in open domains. Currently, there is also no existing dataset that evaluates LLMs on open-domain knowledge that requires tools to solve. To this end, we introduce OpenAct benchmark to evaluate the open-domain task-solving capability, which is built on human expert consultation and repositories in GitHub. It comprises 339 questions spanning 7 diverse domains that need to be solved with domain-specific methods. In our experiments, even state-of-the-art LLMs and LLM-based agents demonstrate unsatisfactory success rates, underscoring the need for a novel approach. Furthermore, we present OpenAgent, a novel LLM-based agent system that can tackle evolving queries in open domains through autonomously integrating specialized tools from GitHub. OpenAgent employs 1) a hierarchical framework where specialized agents handle specific
原文 arXiv:2312.17294;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2312.17294v3