当你的智能体打开聊天应用时:智能体控制的原始聊天日志搜索可与结构化记忆相媲美
When Your Agent Opens the Chat App: Agent-Controlled Search over Raw Chat Logs Rivals Structured Memory
- University of Science and Technology of China(中国科学技术大学)
- MetaStone Technology(MetaStone科技公司)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究提出无语义结构的智能体控制搜索界面ReFind,在MemoryAgentBench等任务上,其检索准确率优于HippoRAG 2等结构化记忆系统,证明可控词汇检索可替代复杂记忆结构实现高效会话记忆任务。
中文摘要 AI 辅助
智能体记忆系统日益通过结构来换取检索质量,在任何问题提出前,就将原始对话历史转化为摘要、嵌入、树或知识图谱。本文探究这种收益有多少来自结构本身,而非对原始历史的有效检索。我们提出ReFind,一种完全不构建语义结构的智能体控制搜索界面:它保留对话档案不做修改,以轮次粒度进行词汇索引,并结合通用迭代关键词搜索循环与四项基于经验重寻工作的聊天原生控制机制:会话感知排名融合、局部上下文扩展、时间范围缩小、跳过已检查会话。单独的推理阶段会基于收集到的证据给出答案。在涵盖会话记忆任务的广泛套件中(单跳与多跳问答、事件排序、事实整合),共约2800个针对精确检索与事实追踪能力的问题,在MemoryAgentBench的增量多轮设置下评估,ReFind达到了所有对比系统中最高的平均准确率(58.2),优于最强的图和树基记忆系统(HippoRAG 2,53.2),所有系统均使用与所有复用基线匹配的GPT-4o-mini主干模型。与单次BM25、匹配的通用智能体BM25对照、组件移除、智能体密集/混合变体的受控对比,分别验证了智能体控制、聊天原生控制、词汇检索的作用。在LongMemEval-S/M上,同一界面使用GPT-5-mini时达到93.2±3.3和89.3±6.0。结果表明,对于针对聊天档案的精确、基于证据的问题,许多归功于复杂记忆结构的收益,可通过让智能体对未修改记录进行可控搜索来实现,完全无需基于LLM的索引构建。
英文摘要
Agent-memory systems increasingly buy retrieval quality with structure, transforming raw conversation histories into summaries, embeddings, trees, or knowledge graphs before any question is asked. We ask how much of that benefit comes from the structure itself, rather than from competent retrieval over the raw history. We present ReFind, an agent-controlled search interface that builds no semantic structure at all: it leaves the conversation archive unmodified, indexes it lexically at turn granularity, and combines a generic iterative keyword-search loop with four chat-native controls grounded in empirical refinding work: session-aware rank fusion, local context expansion, temporal narrowing, and skipping already-inspected sessions. A separate reasoning stage answers from the collected evidence. Across a broad suite of conversational-memory tasks (single- and multi-hop QA, event ordering, and fact consolidation), roughly 2,800 questions on precise-retrieval and fact-tracking capabilities evaluated under the incremental multi-turn setting of MemoryAgentBench, ReFind attains the highest mean accuracy (58.2) of any system compared, above the strongest graph- and tree-based memory systems (HippoRAG 2, 53.2), all under a GPT-4o-mini backbone matched to every reused baseline. Controlled comparisons to single-shot BM25, a matched generic-agentic BM25 control, component removals, and agentic dense/hybrid variants separately support the roles of agent control, chat-native controls, and lexical retrieval. On LongMemEval-S/M, the same interface reaches 93.2 +/- 3.3 and 89.3 +/- 6.0 with GPT-5-mini. The results indicate that for precise, evidence-grounded questions over chat archives, much of the benefit credited to elaborate memory structures is recoverable by giving an agent controllable search over the unmodified record, with no LLM-based index construction at all.