Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow
Wenqi Zhang Affiliation: College of Computer Science and Technology, Zhejiang University{zhangwenqi, Page: https://github.com/zwq2018/Data-Copilot Yongliang Shen Affiliation: College of Computer Science and Technology, Zhejiang University{zhangwenqi, Page: https://github.com/zwq2018/Data-Copilot Zeqi Tan Affiliation: College of Computer Science and Technology, Zhejiang University{zhangwenqi, Page: https://github.com/zwq2018/Data-Copilot Guiyang Hou Affiliation: College of Computer Science and Technology, Zhejiang University{zhangwenqi, Page: https://github.com/zwq2018/Data-Copilot Weiming Lu, Yueting Zhuang Affiliation: College of Computer Science and Technology, Zhejiang University{zhangwenqi, Page: https://github.com/zwq2018/Data-Copilot Affiliation: College of Computer Science and Technology, Zhejiang University{zhangwenqi, Page: https://github.com/zwq2018/Data-Copilot
Abstract
Industries such as finance, meteorology, and energy generate vast amounts of heterogeneous data daily. Efficiently managing, processing, and visualizing such data is labor-intensive and frequently necessitates specialized expertise. Leveraging large language models (LLMs) to develop an automated workflow presents a highly promising solution. However, LLMs are not adept at handling complex numerical computations and table manipulations, and they are further constrained by a limited length context. To bridge this, we propose Data-Copilot, a data analysis agent that autonomously performs data querying, processing, and visualization tailored to diverse human requests. The advancements are twofold: First, it is a code-centric agent that leverages code as an intermediary to process and visualize massive data based on human requests, achieving automated large-scale data analysis. Second, Data-Copilot involves a data exploration phase in advance, which autonomously explores how to design universal and error-free interfaces from data, reducing the error rate in real-time responses. Specifically, It imitates common requests from data sources, abstracts them into universal interfaces (code mo
原文 arXiv:2306.07209;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2306.07209v8