ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models
Haoran Luo, Haihong E , Zichen Tang, Shiyao Peng, Yikai Guo, Wentai Zhang, Thanks: Corresponding author. Affiliation: School of Computer Science, Beijing University of Posts and Telecommunications, China Affiliation: School of Computer Science, Beijing University of Posts and Telecommunications, China Affiliation: School of Computer Science, Beijing University of Posts and Telecommunications, China Affiliation: School of Computer Science, Beijing University of Posts and Telecommunications, China Affiliation: School of Computer Science, Beijing University of Posts and Telecommunications, China Chenghao Ma, Guanting Dong, Meina Song, Wei Lin, Yifan Zhu, Luu Anh Tuan Affiliation: School of Computer Science, Beijing University of Posts and Telecommunications, China Affiliation: School of Computer Science, Beijing University of Posts and Telecommunications, China Affiliation: School of Computer Science, Beijing University of Posts and Telecommunications, China Affiliation: School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China Affiliation: Beijing Institute of Computer Technology and Application Inspur Group Co., Ltd., China Affiliation: College of Computing and Data Science, Nanyang Technological University, Singapore{luohaoran, ehaihong,
Abstract
Knowledge Base Question Answering (KBQA) aims to answer natural language questions over large-scale knowledge bases (KBs), which can be summarized into two crucial steps: knowledge retrieval and semantic parsing. However, three core challenges remain: inefficient knowledge retrieval, mistakes of retrieval adversely impacting semantic parsing, and the complexity of previous KBQA methods. To tackle these challenges, we introduce ChatKBQA, a novel and simple generate-then-retrieve KBQA framework, which proposes first generating the logical form with fine-tuned LLMs, then retrieving and replacing entities and relations with an unsupervised retrieval method, to improve both generation and retrieval more directly. Experimental results show that ChatKBQA achieves new state-of-the-art performance on standard KBQA datasets, WebQSP, and CWQ. This work can also be regarded as a new paradigm for combining LLMs with knowledge graphs (KGs) for interpretable and knowledge-required question answering. Our code is publicly available11 1 https://github.com/LHRLAB/ChatKBQA.
原文 arXiv:2310.08975;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2310.08975v3