arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33013cs.AI

智能体记忆的认识论:长时程LLM智能体中巩固决策的度量与治理

The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents

Sasank Annapureddy, Anjaneya Prasad Thamatani

AI总结:

本研究通过四阶段计划度量并治理长时程LLM智能体的记忆巩固决策,从记忆量转向决策质量与可信度,提出ConsolidationBench基准和受治理的巩固机制,发现质量分数无法预测真实迁移。

AI中文摘要:

长时程LLM智能体必须将积累的经验转化为持久记忆,决定保留什么、压缩什么、抽象为可复用的技能和规则,或遗忘什么。我们报告了一项关于这一巩固问题的四阶段研究计划,其核心发现是度量对象的转变:从智能体记住多少,到其巩固决策是否良好,再到这些决策是否可信。第一阶段通过下游效用从智能体轨迹中学习情节边界;这是一个诚实的近似命中(预言机相关性0.691,对比0.70的基准线),其持久产出是一个三闸防泄漏协议。第二阶段学习在令牌预算下何时提升经验以及提升到何种抽象级别,实现了经验证的+22.7%任务成功率提升和7倍压缩,但暴露了退化遗忘失败和一种我们命名为lambda-流行度耦合的分布偏移失败模式。第三阶段引入ConsolidationBench,一个按构造为预言机的基准,在三个非循环轴上将巩固决策与已知最优解进行评分;生产检索系统保留信息但在跨级别迁移上得分为零。第四阶段引入受治理的巩固:将决策包裹在抗毒、可逆性和可审计性保证以及质量门中。治理在统计上与质量分数不同(r²=0.43;偏相关r=0.27;相同质量的政策在治理上相差三倍),因此该贡献独立于指标的外部有效性而存在。关于该问题,我们报告了一个已解决的否定结果:在分级重用重新设计移除结构性上限后,一项包含2,532个真实答案单元的双基准研究发现质量分数不能预测真实迁移准确性(合并Spearman ρ=-0.24,n=12,置信区间跨越零)。一次对抗性自我批判过程清除了最终声明集,零幸存过度声明。

英文摘要:

Long-horizon LLM agents must convert accumulated experience into durable memory, deciding what to keep, compress, abstract into reusable skills and rules, or forget. We report a four-phase research program on this consolidation problem whose central finding is a shift in what is measured: from how much an agent remembers, to whether its consolidation decisions are any good, to whether those decisions can be trusted. Phase 1 learns episodic boundaries from agent traces by downstream utility; an honest near-miss (oracle correlation 0.691 vs a 0.70 bar) whose lasting output is a three-gate anti-leakage protocol. Phase 2 learns when to promote experience and to which abstraction level under a token budget, achieving a verified +22.7% task-success improvement with 7x compression, but exposing a degenerate-forgetting failure and a distribution-shift failure mode we name lambda-prevalence coupling. Phase 3 introduces ConsolidationBench, an oracle-by-construction benchmark that scores consolidation decisions against a known optimum on three non-circular axes; production retrieval systems retain information yet score zero on cross-level transfer. Phase 4 introduces governed consolidation: the decision wrapped in poison-resistance, reversibility, and auditability guarantees with a quality gate. Governance is statistically distinct from the quality score ($r^2 = 0.43$; partial $r = 0.27$; identical-quality policies differ threefold in governance), so the contribution survives independently of the metric's external validity. On that question we report a resolved negative: after a graded-reuse redesign removed a structural ceiling, a two-benchmark study with 2,532 real answer cells finds the quality score does not predict real transfer accuracy (pooled Spearman $ρ= -0.24$, n = 12, CI spanning zero). An adversarial self-critique pass cleared the final claim set with zero surviving overclaims.

补充信息

↑