Evaluating Verifiability in Generative Search Engines
Nelson F. Liu Tianyi Zhang Percy Liang Computer Science Department Stanford University
Abstract
Generative search engines directly generate responses to user queries, along with in-line citations. A prerequisite trait of a trustworthy generative search engine is verifiability, i.e., systems should cite comprehensively (high citation recall; all statements are fully supported by citations) and accurately (high citation precision; every cite supports its associated statement). We conduct human evaluation to audit four popular generative search engines—Bing Chat, NeevaAI, perplexity.ai, and YouChat—across a diverse set of queries from a variety of sources (e.g., historical Google user queries, dynamically-collected open-ended questions on Reddit, etc.). We find that responses from existing generative search engines are fluent and appear informative, but frequently contain unsupported statements and inaccurate citations: on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence. We believe that these results are concerningly low for systems that may serve as a primary tool for information-seeking users, especially given their facade of trustworthiness. We hope that our results further motivate th
中文速览
生成式搜索引擎(generative search engine)会直接对用户提问给出完整答案并附带引用来源,但这些引用到底靠不靠谱,此前缺乏系统性评估。研究者对Bing Chat、NeevaAI、perplexity.ai和YouChat四款主流产品进行了人工审计,从引用召回率(citation recall,即有多少陈述被引文充分支持)和引用精确率(citation precision,即有多少引文真正支撑了对应陈述)两个维度衡量其可验证性。结果令人担忧:平均只有51.5%的生成句子被引文充分支持,而74.5%的引文才真正支撑了其关联陈述;更糟糕的是,看起来越有用的回答,其引用精确率反而越低,说明这类系统往往靠直接摘抄网页内容来制造"信息丰富"的假象。这一发现揭示了生成式搜索引擎在可信度上的严重缺口,对于把它当作主要信息来源的普通用户而言风险不容忽视,也呼吁学界和业界加快建立更可靠的评测标准与系统设计规范。
原文 arXiv:2304.09848;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2304.09848v2