发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有LLM多智能体系统记忆系统无法建模智能体可信度的问题,提出Σ-Mem在线可靠性记忆,基于决策后反馈更新实对称状态,在Qwen系列模型上实现自适应协调与更优性能。
AI 中文摘要
记忆是长程LLM智能体的核心,但现有记忆系统主要存储交互内容,而非建模哪些智能体可信及在何种条件下可信。这一局限在多智能体系统中尤为突出,因为中心模型可能无法直接验证合理或相关的同伴响应。我们提出Σ-Mem,一种在线可靠性记忆,它记录单个同伴的历史能力证据,以及同伴集合内的同伴关系证据。两种形式的证据均以实对称状态维护,并根据决策后的正确性反馈进行更新。根据魏尔不等式,每个事件级更新引起的谱变化是有界的,从而无需对底层模型进行重新训练即可实现稳定的在线自适应。Σ-Mem提供通用的读写接口:同一记忆可用于中心模型的残差引导、无响应同伴路由或可靠性加权投票。在五个通义千问(Qwen)系列模型上,Σ-Mem可适应反事实可靠性变化,并能泛化到未见过的同伴和任务领域。在全部分布外(OOD)评估集上,直接读取记忆的表现优于多数投票和表现最佳的固定同伴。此外,随着更多正确性反馈的获取,性能持续提升,表明Σ-Mem逐步积累可操作的可靠性信息。这些结果确立了可靠性记忆作为基于LLM的多智能体系统自适应协调的可复用基础的地位。
英文摘要
Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rather than modeling which agents can be trusted and under what conditions. This limitation is particularly important in multi-agent systems, where a central model may be unable to directly verify plausible or correlated peer responses. We introduce $Σ$-Mem, an online reliability memory that records historical competence evidence for individual peers and peer relationship evidence across the peer set. Both forms of evidence are maintained as real symmetric states and updated from post-decision correctness feedback. By Weyl's inequality, the spectral change caused by each event-level update is bounded, enabling stable online adaptation without retraining the underlying models. $Σ$-Mem provides a general write-and-read interface: the same memory can be used for residual steering of a central model, response-free peer routing, or reliability-weighted voting. Across five Qwen-family models, $Σ$-Mem adapts to counterfactual reliability shifts and generalizes to unseen peers and task domains. Direct memory readouts also outperform majority voting and the best fixed peer over the full OOD evaluation set. Moreover, performance improves consistently as more correctness feedback becomes available, indicating that $Σ$-Mem progressively accumulates actionable reliability information. These results establish reliability memory as a reusable foundation for adaptive coordination in LLM-based multi-agent systems.