GPT4Graph: Can Large Language Models Understand Graph Structured Data? An Empirical Evaluation and Benchmarking
Jiayan Guo11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPT111, Lun Du22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT222Corresponding Author, Hengyu Liu33{}^{3}start_FLOATSUPERSCRIPT 3 end_FLOATSUPERSCRIPT, Mengyu Zhou22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT, Xinyi He44{}^{4}start_FLOATSUPERSCRIPT 4 end_FLOATSUPERSCRIPT, Shi Han22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPT 11{}^{1}start_FLOATSUPERSCRIPT 1 end_FLOATSUPERSCRIPTSchool of Intelligence Science and Technology, Peking University; 22{}^{2}start_FLOATSUPERSCRIPT 2 end_FLOATSUPERSCRIPTMicrosoft; 33{}^{3}start_FLOATSUPERSCRIPT 3 end_FLOATSUPERSCRIPTUniversity of Technology Sydney; 44{}^{4}start_FLOATSUPERSCRIPT 4 end_FLOATSUPERSCRIPTXi’an Jiaotong University
Abstract
Large language models (LLM) like ChatGPT have become indispensable to artificial general intelligence (AGI), demonstrating excellent performance in various natural language processing tasks. Graph data is ubiquitous and an essential part of AGI. The training corpus of large language models often includes some algorithmic components, which allows them to achieve certain effects on some graph data-related problems. However, there is still little research on their performance on a broader range of graph-structured data. In this paper, we conduct an empirical study to assess the proficiency of LLMs in comprehending graph data, employing a diverse range of structural and semantic-related tasks that evaluate the LLMs’ capabilities in graph understanding. Through our study, we uncover current limitations and future directions of LLMs in comprehending graph and performing associated reasoning tasks.
中文速览
大型语言模型(LLM)如ChatGPT在文本任务上表现出色,但它们究竟能不能真正"读懂"图结构数据(graph-structured data),目前研究还很匮乏。为此,这篇论文构建了一套系统性的评测框架,将图数据转化为模型可读的文本描述语言,并设计了涵盖结构理解(如节点度数检测、图直径计算)和语义理解(如知识图谱问答、节点分类)共十类任务的基准测试,同时对比了手工提示和模型自生成提示等多种提示策略。实验结果显示,LLM在简单的图结构感知任务上有一定能力,但在需要多步推理的复杂结构计算以及语义理解任务上,与专门针对图设计的模型相比仍有显著差距。这项工作为厘清LLM的图理解边界、指引后续改进方向提供了重要参考,对推动语言模型迈向真正的通用人工智能(AGI)具有现实意义。
原文 arXiv:2305.15066;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2305.15066v2