超越记忆构建:重新思考基于LLM的对话代理的记忆访问
Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents
AI总结:
本文针对长时程高熵对话中基于LLM的记忆构建易损且成本高的问题,提出Threader系统,通过保留原始交互并采用结构感知的多信号检索,显著提升答案准确性与证据召回率,同时降低构建开销。
AI中文摘要:
记忆是对话代理的核心组成部分,能够在长时间交互中实现连贯且上下文感知的行为。近期方法通常依赖于基于LLM的记忆构建,即将原始交互重写为结构化的记忆单元,并通过RAG管道进行后续检索。尽管在受控环境中有效,但我们表明,在长时程、高熵对话中,这种范式会失效:随着上下文长度和信息复杂度的增长,记忆构建变得越来越有损且不稳定,并且由于反复调用LLM而产生高昂成本。为解决这些局限性,我们提出了Threader,一种将焦点从记忆构建转向对原始交互的高效、结构感知访问的记忆系统。Threader不重写交互,而是将其保留为一级记忆,通过轻量级增量分段将其组织为主题连贯的片段,并通过多视图表示实现精确检索。在查询时,它执行多信号检索,结合片段级访问与局部证据匹配,确保完整性和连贯性。大量实验表明,Threader持续提高了答案准确性和证据召回率,同时显著降低了记忆构建开销。
英文摘要:
Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and later retrieved via a RAG pipeline. While effective in controlled settings, we show that this paradigm breaks down in long-horizon, high-entropy conversations: memory construction becomes increasingly lossy and unstable as context length and information complexity grow, and incurs prohibitive cost due to repeated LLM invocation. To address these limitations, we propose Threader, a memory system that shifts the focus from memory construction to efficient, structure-aware access over raw interactions. Instead of rewriting interactions, Threader preserves them as first-class memory, organizes them into topic-coherent segments via lightweight incremental segmentation, and enables accurate retrieval through multi-view representation. At query time, it performs multi-signal retrieval that combines segment-level access with localized evidence matching, ensuring both completeness and coherence. Extensive experiments demonstrate that Threader consistently improves answer accuracy and evidence recall, while significantly reducing the memory construction overhead.