arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23465cs.CLcs.AI

提出、验证、提交:面向长时程多智能体对话的基于证据的记忆机制

Propose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations

  • Alibaba Cloud Computing(阿里云计算)

机构由 AI 辅助整理,请以论文原文为准。

Zihao Lu, Zhihang Yuan, Lei Shi

AI总结:

针对长时程多智能体对话记忆挑战,提出基于证据的提出-验证-提交协议和可搜索状态机EGMEMORY,在多个基准上显著超越基线,并展示良好泛化性。

AI中文摘要:

长时程对话记忆在多智能体场景中尤为具有挑战性,因为相关证据分散在不同参与者和上下文中,且先前建立的信息可能随后被修订。我们提出EGMEMORY,将长时程多智能体记忆构建为一个可搜索的状态机,将持久的消息级证据与显式的活动状态分离。在写入时,自适应状态解析和基于证据的提出-验证-提交协议控制该状态的演化。在读取时,自适应证据导航迭代地解析查询所需的状态和支持证据,利用对话结构缩小搜索空间,并利用词汇-语义相关性对候选进行排序。该系统通过提示和工具使用运行,无需针对记忆的策略训练。EGMEMORY在GroupMemBench上达到68.2%,在EverMemBench上达到77.9%,分别比最强评估基线高出22.7和21.4个百分点。它进一步在双人对话基准LoCoMo上达到73.6%,展示了超越多智能体对话的泛化能力。我们将在正式发表后发布代码库。

英文摘要:

Long-horizon conversational memory is especially challenging in multi-actor settings, where relevant evidence is distributed across participants and contexts and previously established information may later be revised. We introduce EGMEMORY, which formulates long-horizon multi-actor memory as a searchable state machine that separates persistent message-level evidence from an explicit active state. At write time, adaptive state resolution and an evidence-grounded propose-verify-commit protocol govern how this state evolves. At read time, adaptive evidence navigation iteratively resolves the state and supporting evidence required for a query, using conversational structure to narrow the search space and lexical-semantic relevance to rank candidates. The system operates through prompting and tool use without memory-specific policy training. EGMEMORY achieves 68.2% on GroupMemBench and 77.9% on EverMemBench, outperforming the strongest evaluated baselines by 22.7 and 21.4 percentage points, respectively. It further reaches 73.6% on the dyadic LoCoMo benchmark, demonstrating generalization beyond multi-actor conversations. We will release the codebase upon formal publication.

↑