arXiv:2508.13171 · 中英对照阅读
认知工作空间:面向大语言模型的主动记忆管理——功能性无限上下文的实证研究
Cognitive Workspace: Active Memory Management for LLMs An Empirical Study of Functional Infinite Context
中文速览
长上下文和 RAG 虽能让大语言模型接触更多信息,却仍缺少像人一样主动整理、保留和调用记忆的能力,导致上下文管理低效且缺乏连续性。论文提出“认知工作空间”(Cognitive Workspace),借鉴工作记忆、外部认知和分布式认知理论,通过主动筛选信息、分层保存工作状态,并根据任务需求动态优化上下文,把外部记忆变成思考过程的一部分。实验显示,该方法的记忆复用率平均达到58.6%,传统 RAG 为0%,即使操作次数多3.3倍,整体效率仍提升17%–18%,且统计结果显著。其重要性在于,它把大模型的记忆设计从“查资料”推进到“主动管理认知状态”,为处理超长任务、跨会话推理和更稳定的智能体系统提供了新方向。
摘要
Large Language Models (LLMs) face fundamental limitations in context management despite recent advances extending context windows to millions of tokens. We propose Cognitive Workspace, a novel paradigm that transcends traditional Retrieval-Augmented Generation (RAG) by emulating human cognitive mechanisms of external memory use. Drawing from cognitive science foundations including Baddeley’s working memory model [1, 2], Clark’s extended mind thesis [3], and Hutchins’ distributed cognition framework [4], we demonstrate that current passive retrieval systems fail to capture the dynamic, task-driven nature of human memory management. Our analysis of 2024-2025 developments reveals that while techniques like Infini-attention [5] and StreamingLLM [6] achieve impressive context lengths, they lack the metacognitive awareness and active planning capabilities essential for true cognitive extension. Cognitive Workspace addresses these limitations through three core innovations: (1) active memory management with deliberate information curation, (2) hierarchical cognitive buffers enabling persistent working states, and (3) task-driven context optimization that dynamically adapts to cognitive demands.
术语表
- Large Language Models (LLMs)
- 大语言模型(LLM)
- Cognitive Workspace
- 认知工作空间
- Retrieval-Augmented Generation (RAG)
- 检索增强生成(RAG)
- working memory
- 工作记忆
- extended mind thesis
- 延展心智论
- distributed cognition
- 分布式认知
- Infini-attention
- Infini-attention
- StreamingLLM
- StreamingLLM
- metacognitive awareness
- 元认知意识
- active memory management
- 主动记忆管理
- hierarchical cognitive buffers
- 分层认知缓冲区
- task-driven context optimization
- 任务驱动的上下文优化
- memory reuse rate
- 记忆复用率
- functional infinite context
- 功能性无限上下文
- Transformer architecture
- Transformer架构
- quadratic attention complexity
- 二次注意力复杂度
- FlashAttention 3
- FlashAttention 3
- MInference
- MInference
- MemGPT
- MemGPT
- Hierarchical Memory Transformer
- 分层记忆Transformer
- Self-RAG
- Self-RAG
- CRAG
- CRAG
- Adaptive RAG
- Adaptive RAG
- cognitive artifacts
- 认知人工制品
- Baddeley’s multicomponent working memory model
- 巴德利多成分工作记忆模型
- central executive
- 中央执行系统
- phonological loop
- 语音回路
- visuospatial sketchpad
- visuospatial sketchpad
- episodic buffer
- 情景缓冲器
- Cowan’s embedded-processes model
- 考恩嵌入过程模型