arXiv:2601.00821 · 中英对照阅读
保真度优先于结构:在长对话 LLM 记忆中,逐字对话块胜过有损结构化工件提取
Fidelity Before Structure: Verbatim Chunks Beat Lossy Artifact Extraction in Long-Conversation LLM Memory
中文速览
长对话记忆系统常把原文压缩成事实、决定等结构化条目,但关键问题是这种“提炼”会不会丢掉回答问题所需的细节。研究者在同一套检索、重排和推理流程中,只替换存储形式,严格比较原始对话片段与大模型抽取的结构化记忆,并用语义图和多项控制实验排除其他因素。结果显示,原文片段在 LoCoMo 上领先 15.9 个百分点、在 LongMemEval-S 上领先 22 个百分点,语义图也无法弥补差距,说明主要损失来自有损压缩而非缺少结构。这个结论提醒我们,结构化记忆更适合与原文并存,而不应直接取代原文;在当前这类无训练、无状态的系统里,保真度比“看起来更聪明”的抽象结构更重要。
摘要
A growing class of conversational-memory systems compresses dialogue history into structured artifacts (extracted facts, decisions, or events) on the premise that distilled structure retrieves better than raw text. We test this premise with a controlled ablation: within one fixed retrieval–rerank–reasoning pipeline, we swap only the stored representation (LLM-extracted typed artifacts versus verbatim conversation chunks), holding the model, retriever, reranker, and judge constant. Verbatim chunks win by 15.9 points on LoCoMo (43.9% vs. 28.0%) and 22.0 points on LongMemEval-S (67.4% vs. 45.4%); a 1-hop semantic graph does not recover the gap, and six confound controls reproduce the effect. The mechanism is lossy distillation, not structure per se: accuracy tracks how much source text survives in the store, and the extracted-artifact pipeline does not beat naive RAG in overall accuracy (though chunks abstain worse, §6). For the extraction designs we test, structured memory should augment verbatim text rather than replace it: adding artifacts alongside chunks preserves accuracy; substituting them forfeits the gap. Code and data: https://github.com/tao-hpu/cog-canvas.
术语表
- conversational-memory systems
- 对话记忆系统
- structured artifacts
- 结构化工件
- extracted facts
- 抽取事实
- retrieval–rerank–reasoning pipeline
- 检索–重排序–推理流水线
- verbatim conversation chunks
- 逐字对话块
- retriever
- 检索器
- reranker
- 重排序器
- LLM judge
- LLM 评判器
- LoCoMo
- LoCoMo
- LongMemEval-S
- LongMemEval-S
- 1-hop semantic graph
- 一跳语义图
- lossy distillation
- 有损蒸馏
- naive RAG
- 朴素 RAG
- retrieval-augmented generation (RAG)
- 检索增强生成(RAG)
- typed artifacts
- 带类型工件
- structured memory
- 结构化记忆
- full-context baseline
- 完整上下文基线
- exact match
- 精确匹配
- fidelity
- 保真度
- ablation study
- 消融实验
- confound control
- 混杂因素控制
- GraphRAG
- GraphRAG
- dense-only retrieval
- 仅密集检索
- source provenance
- 源出处
- memory-augmented LLMs
- 记忆增强型 LLM
- virtual memory management
- 虚拟内存管理
- stateless, training-free regime
- 无状态、免训练范式
- semantic graph
- 语义图
- sparse attention
- 稀疏注意力