AI 中文总结
提出LAM有损智能体记忆框架,通过确定性去重规则(带检索分数误差界)、内存管理器与性能模型,在600条轨迹上移除22.47%观察令牌并保留99.984%证据,实现71.4-91.6倍端到端加速。
AI 中文摘要
智能体记忆随着智能体读取输入、推理和调用工具而增长。更长的历史记录会增加推理成本,并最终超出上下文窗口。基于LLM的摘要可以缩减历史记录,但会增加延迟,并且对信息损失没有明确的界限。我们提出了LAM,一种有损智能体记忆系统,包含三个组件:一个确定性的去重规则,带有检索分数的替换界限——该界限是对分数扰动的约束,而非排名不变的保证;一个内存管理器,保留缓存的前缀并将压缩与推理重叠;以及一个性能模型,在部署前估计压缩成本。在600条智能体轨迹上,LAM移除了22.47%的观察令牌,同时保留了99.984%的实测黄金补丁证据。在固定的删除集下,性能模型预测,在预填充前移除记录相比从预填充上下文中删除记录,可实现71.4倍至91.6倍的端到端加速。该收益来自调度而非规则,并适用于任何保留前缀的测试。
英文摘要
Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores - a bound on score perturbation, not a certificate of unchanged ranking; a memory manager that preserves the cached prefix and overlaps compaction with inference; and a performance model that estimates compaction costs before deployment. On 600 agent trajectories, LAM removes 22.47% of observation tokens while retaining 99.984% of the measured gold-patch evidence. At a fixed deletion set, the performance model predicts a 71.4x-91.6x end-to-end speedup from removing records before prefill instead of deleting them from a prefilled context. That benefit comes from the schedule rather than the rule and applies to any prefix-preserving test.