arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

记忆信任缺口:持久记忆智能体中依赖能力的故障

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

Jundong Hu, Shekar Ramachandran

arXiv 2609.01852首次发表:更新:

发表机构

PayPal AI(PayPal人工智能部门)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对Qwen3等模型,发现持久记忆智能体存在依赖能力的记忆信任缺口,过时信息会被过度信任,且缓解措施需适配模型能力,在多模型和数据集上验证了该现象。

AI 中文摘要

持久记忆支持个性化智能体,但过时存储的事实会在无预警的情况下覆盖当前权威证据。我们研究当模型能力变化时,这种危害何时开始显现。我们评估了一个冻结的闭集动作评分基准,该基准包含两组套件,分别代表“无记忆”的两种不同含义:Benefit套件(若无存储事实则无法解决问题)和Safety套件(其中权威工具始终持有正确值),测试对象为同一系列的模型规模(Qwen3 0.6/1.7/4/8B)。记忆信任缺口反映的是过度信任而非混淆。在Benefit套件中,各规模模型以0.92-1.00的概率用过时值作答。在Safety套件中,陷阱条件(Δ_mem)下低于无记忆基线的危害是能力门控的,一旦过时记录被伪装成当前状态,较大模型大多会崩溃。在2×2×2×2的因子设计中,触发过度信任的特征取决于特征和模型规模。移除标签会放大所有规模模型的过度信任,而近期性特征(过时日期更新)对较大模型的欺骗性更强。源权威性较弱且与规模无关,其影响在Qwen3模型规模系列中从正变负。我们通过直接跨规模对比测试而非每个模型的重叠区间来确认这些规模交互作用。缓解措施同样依赖能力:暴露元数据可提升能力较强模型的准确率,但仅预先解决冲突才能恢复2个较小检查点的准确率。该模式在独立的Llama-Instruct模型规模系列的能力较强模型以及2个外部数据集(RGB、MisBench)上也存在。框架控制未发现记忆标签的一致优势:在3个较小规模下,模型更信任过时文档而非过时记忆;在8B规模下,差异不显著。

英文摘要

Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capability changes. We evaluate a frozen, closed-set, action-scored benchmark with 2 suites that represent 2 different meanings of "no memory" (a Benefit suite, unsolvable without the stored fact, and a Safety suite, in which an authoritative tool always holds the correct value), on a same-family model-size series (Qwen3 0.6/1.7/4/8B). The Memory Trust Gap reflects over-trust rather than confusion. In the Benefit suite, models answer with the stale value 0.92-1.00 of the time at every scale. In the Safety suite, harm below the no-memory baseline under the trap conditions ($Δ_{\mathrm{mem}}$) is capability-gated, with the larger models collapsing most once a stale note is made to look current. In a $2\times2\times2\times2$ factorial, which feature triggers over-trust depends on both the feature and model scale. Removing a label amplifies over-trust at every size, and a recency feature (stale dated newer) fools the larger models harder. Source authority is weak and scale-flat, and position changes from positive to negative across the Qwen3 model-size series. We confirm these scale interactions with direct cross-size contrast tests rather than overlapping per-model intervals. Mitigation is likewise capability-dependent: exposing metadata improves accuracy for the capable models, but only pre-resolving the conflict restores accuracy for the 2 smaller checkpoints. The same pattern appears on the capable models in an independent Llama-Instruct model-size series and on 2 external datasets (RGB, MisBench). A framing control finds no consistent advantage for the memory label: at the 3 smaller scales, models trust a stale document more than a stale memory; at 8B, the difference is not significant.

CommentsPreprint. Under review at a NeurIPS 2026 workshop. 14 pages, 7 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑