跨回合智能体记忆的失效契约
Invalidation Contracts for Cross-Episode Agent Memory
- South Dakota State University(南达科他州立大学)
- Siemens Digital Industries Software(西门子数字工业软件)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对LLM智能体缓存API错误修复方案遇数据漂移失效的问题,提出失效契约协议,通过附加版本戳等机制提升合规性与token效率,在多模型多回合实验中验证了其有效性。
AI中文摘要:
缓存API错误恢复建议的大语言模型(LLM)智能体可在后续回合中跳过重新推导,减少在已掌握约束上的token消耗和模型调用次数。服务器端数据漂移会使这些缓存的修复方案变成静默失败,而常规补救措施——每回合都重新推导——会抵消上述节省。我们提出失效契约(invalidation contracts),这是一个协议层,为每条恢复建议附加版本戳和可缓存性提示,使客户端无需反复尝试即可移除过时条目,同时保留其余有效条目。该契约将实际节省分解为两个独立因素:有效性(validity,漂移事件后仍正确的缓存建议比例,仅取决于协议且与供应商无关)和合规性(compliance,规划器首次尝试应用的比例,取决于规划器模型)。相同的网络字节在Claude Haiku 4.5上实现100%首次尝试合规,而在表现出输入模式保守性(拒绝添加原始请求未包含字段的修复方案)的Claude Sonnet 5上仅为11%或更低。我们在7种模型、3种服务路径、2个领域和约9400个回合中进行评估:行级失效使7种模型的合规性提升0至66.7个百分点,其中3种模型提升55.6至66.7个百分点,在7种模型中的4种上恢复了基线token成本的29%-33%;而表级失效会破坏共置条目,使7种模型中的5种的漂移后首次尝试率降至0%。在第4.1节的行级预言机下,行粒度的驱逐精度在所有模型上均为1.00。该契约使响应负载增加15%,版本戳有效性是确定性的,在所有模型和服务路径上产生相同结果,整个评估过程中零契约失败。
英文摘要:
LLM agents that cache recovery suggestions from API errors can skip re-derivation in later episodes, spending fewer tokens and fewer model calls on constraints they have already learned. Server-side data drift turns those cached fixes into silent failures, and the usual remedy, re-deriving on every episode, gives the savings back. We introduce invalidation contracts, a protocol layer that attaches version stamps and cacheability hints to every recovery suggestion so the client can evict stale entries without trial and error, and keep the rest. The contract decomposes realized savings into two independent factors: validity, the fraction of cached suggestions that remain correct after a drift event, and compliance, the fraction the planner applies on the first attempt. Validity depends only on the protocol and is vendor-independent. Compliance depends on the planner model: identical wire bytes yield 100% first-try compliance on Claude Haiku 4.5 and 11% or below on Claude Sonnet 5, which exhibits input-schema conservatism, refusing fixes that add fields the original request did not contain. We evaluate across seven models, three serving paths, two domains, and approximately 9,400 episodes. Row-level invalidation raises compliance by 0 to 66.7 percentage points across the seven models, 55.6 to 66.7 on three, and recovers 29-33% of baseline token cost on four of seven models, while table-level invalidation destroys co-located entries and drops post-drift first-try rates to 0% on five of seven. Eviction precision is 1.00 at row granularity on every model under the row-level oracle of Section 4.1. The contract adds 15% to response payload. Version-stamp validity is deterministic by construction and produced identical results across every model and serving path, with zero contract failures in the entire evaluation.