arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33517cs.MA

TRACE:演化多智能体系统中的记忆有效性治理

TRACE: Governing Memory Validity in Evolving Multi-Agent Systems

  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • UC Berkeley(加州大学伯克利分校)
  • MemTensor(Shanghai) Technology Co., Ltd.(墨通(上海)科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Wenjun Xiong, Shengtao Zhang, Shangding Gu, Bo Tang, Zhiyu Li, Feiyu Xiong, Ying Wen, Muning Wen

AI总结:

针对多智能体系统中返回智能体可能依据过时记忆行动的问题,提出无训练层TRACE,将重新进入视为资格决策,在ManBench-Return上同时实现高有效信息可用性与高无效信息拒绝率。

AI中文摘要:

持久记忆使语言模型智能体能够在长期协作中携带信息,但留下了一个生命周期问题:一旦共享状态发生变化,返回的智能体还可以依据哪些信息采取行动?一段记忆可以被正确检索、与当前任务相关、忠实于其来源,但仍然不可用于行动:暂停前保存的行程仍然提到团队已经更换的酒店。我们将此形式化为时间性记忆准入问题,并提出TRACE,一个无需训练的层,将重新进入视为资格决策而非存储或检索操作,将离开时的检查点与缺席期间的更新进行协调,解决显式和隐式失效,并且仅在返回视图覆盖返回角色的未完成义务时释放有界返回视图。我们在Memora、STALE Type II和派生的ManBench-Return设置下,在三种行为者模型上评估TRACE,每种都重铸为返回情节:一个智能体离开,四个队友改变共享状态,然后智能体重新加入。区分方法的关键不是整体准确性,而是能否同时保留有效记忆和拒绝过时记忆,而没有任何单一策略基线能够做到:Restore(完全恢复离开时的检查点)接受过时状态,Reset(从空记忆开始返回)丢弃有效状态,各自在两项指标之一上降至0%。TRACE是唯一在两者上都表现良好的方法,在ManBench-Return上达到92.6-98.3%的有效信息可用性和98.4-99.5%的无效信息拒绝率,整体准确性在最佳基线的3.8个百分点以内。在STALE Type II上,它相对于最强比较策略在Overall上提高了22.3(Qwen)、18.5(Gemini)和27.5(DeepSeek)个百分点,同时使用约2.3倍于它们的令牌,而写入时整合流水线在TRACE的3.99倍令牌下更加准确。

英文摘要:

Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, and faithful to its source, and nonetheless be inadmissible for action: an itinerary saved before a pause still names the hotel the team has since replaced. We formalize this as temporal memory admission and present TRACE, a training-free layer that treats re-entry as an eligibility decision rather than a storage or retrieval operation, reconciling a departure checkpoint against absence-period updates, resolving explicit and implicit invalidation, and releasing a bounded Return View only when it covers the returning role's open obligations. We evaluate TRACE under three actor models on Memora, STALE Type II, and a derived ManBench-Return setting, each recast as return episodes: one agent departs, four teammates change the shared state, and the agent rejoins. What separates methods is not overall accuracy but whether one can retain valid memory and reject stale memory at once, and no single-policy baseline can: Restore (reinstate the departure checkpoint in full) admits stale state, Reset (start the return from an empty memory) discards valid state, each bottoming out at 0% on one of the two. TRACE is the only method high on both, reaching 92.6-98.3% valid-information availability with 98.4-99.5% invalid-information rejection on ManBench-Return, within 3.8 points of the best baseline's overall accuracy. On STALE Type II it improves Overall over the strongest comparison policy by 22.3 (Qwen), 18.5 (Gemini), and 27.5 (DeepSeek) points at roughly 2.3 times their tokens, while a write-time consolidation pipeline is more accurate still at 3.99 times TRACE's.

补充信息

↑