LLM4DyG: Can Large Language Models Solve Spatial-Temporal Problems on Dynamic Graphs?
Zeyang Zhang 0000-0003-1329-1313 DCST, Tsinghua UniversityBeijingChina , Xin Wang 0000-0002-0351-2939 DCST, BNRist, Tsinghua UniversityBeijingChina , Ziwei Zhang 0000-0003-2451-843X DCST, Tsinghua UniversityBeijingChina , Haoyang Li 0000-0003-3544-5563 DCST, Tsinghua UniversityBeijingChina , Yijian Qin 0000-0002-0419-5226 DCST, Tsinghua UniversityBeijingChina and Wenwu Zhu 0000-0003-2236-9290 DCST, BNRist, Tsinghua UniversityBeijingChina
Abstract
In an era marked by the increasing adoption of Large Language Models (LLMs) for various tasks, there is a growing focus on exploring LLMs’ capabilities in handling web data, particularly graph data. Dynamic graphs, which capture temporal network evolution patterns, are ubiquitous in real-world web data. Evaluating LLMs’ competence in understanding spatial-temporal information on dynamic graphs is essential for their adoption in web applications, which remains unexplored in the literature. In this paper, we bridge the gap via proposing to evaluate LLMs’ spatial-temporal understanding abilities on dynamic graphs, to the best of our knowledge, for the first time. Specifically, we propose the LLM4DyG benchmark, which includes nine specially designed tasks considering the capability evaluation of LLMs from both temporal and spatial dimensions. Then, we conduct extensive experiments to analyze the impacts of different data generators, data statistics, prompting techniques, and LLMs on the model performance. Finally, we propose Disentangled Spatial-Temporal Thoughts (DST2) for LLMs on dynamic graphs to enhance LLMs’ spatial-temporal understanding abilities. Our main observations are: 1) L
中文速览
大语言模型(LLM)能否真正理解动态图(dynamic graph)上的时空信息,此前从未被系统评估过。为此,研究者提出了 LLM4DyG 基准,围绕时间和空间两个维度设计了九项任务,涵盖从"某条边何时出现"到"动态三元闭包是否形成"等不同难度的问题,并在多种数据生成方式、图规模与密度参数、提示策略以及五种 LLM 上展开大规模实验。结果表明,LLM 具备初步的时空推理能力,但随着图规模和密度增大,性能显著下降,而时间跨度和数据生成机制的影响则相对有限;研究者进一步提出"解耦时空思维"(Disentangled Spatial-Temporal Thoughts,DST²)提示方法,引导模型先处理时间信息再处理结构信息,在多数任务上带来明显提升(如"边何时出现"任务准确率从 33.7% 跃升至 76.7%)。这项工作填补了 LLM 在动态图理解方面的评估空白,对序列推荐、趋势预测、欺诈检测等真实网络应用具有重要参考价值。
原文 arXiv:2310.17110;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2310.17110v3