A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
Lei Huang Harbin Institute of Technology800 Dongchuan RoadHarbinHeilongjiangChina150001 , Weijiang Yu Huawei Inc.Bantian SubdistrictShenzhenGuangdongChina518129 , Weitao Ma , Weihong Zhong Harbin Institute of Technology800 Dongchuan RoadHarbinHeilongjiangChina150001 , Zhangyin Feng , Haotian Wang Harbin Institute of Technology800 Dongchuan RoadHarbinHeilongjiangChina150001 , Qianglong Chen , Weihua Peng Huawei Inc.Bantian SubdistrictShenzhenGuangdongChina518129 , Xiaocheng Feng , Bing Qin and Ting Liu Harbin Institute of Technology800 Dongchuan RoadHarbinHeilongjiangChina150001
Abstract
The emergence of large language models (LLMs) has marked a significant breakthrough in natural language processing (NLP), fueling a paradigm shift in information acquisition. Nevertheless, LLMs are prone to hallucination, generating plausible yet nonfactual content. This phenomenon raises significant concerns over the reliability of LLMs in real-world information retrieval (IR) systems and has attracted intensive research to detect and mitigate such hallucinations. Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in c
中文速览
大型语言模型(LLM)虽然在自然语言处理领域取得了巨大突破,却存在一个棘手的"幻觉"(hallucination)问题——它们会生成听起来合理、实则不符合事实的内容,严重威胁聊天机器人、搜索引擎等真实信息系统的可靠性。这篇综述针对LLM时代的幻觉现象提出了新的分类框架,将其分为"事实性幻觉"(factuality hallucination,即与可验证事实不符)和"忠实性幻觉"(faithfulness hallucination,即与用户指令、上下文或自身逻辑不一致),并从数据、训练、推理三个层面系统梳理了幻觉的成因。在此基础上,文章全面综述了现有的幻觉检测方法与评测基准,介绍了涵盖数据处理、模型训练和推理优化的多种缓解策略,还专门分析了当前检索增强生成(RAG)系统在对抗幻觉时面临的固有局限。这项工作为研究者理解和解决LLM幻觉问题提供了迄今最系统的参考框架,对推动更可信赖的AI信息系统的开发具有重要指导价值。
原文 arXiv:2311.05232;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2311.05232v2