发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过LIMBO沙箱实验探究LLM智能体中精确一次执行的归属,发现模型在可回读时起决定作用,契约在不可回读时主导,幂等键能显著降低重复率。
AI 中文摘要
当使用工具的智能体的写入操作超时或返回服务器错误时,该操作可能已经生效。盲目重试会重复执行该操作——导致第二次扣费、第二次公告、第二次部署——而放弃则会跳过必要的工作。我们探究精确一次行为应在何处强制执行:在模型中、在智能体执行框架中,还是在工具契约中。我们引入了LIMBO,一个由六个服务组成的确定性沙箱,具有现实的契约(可选的幂等键、最终一致且缺失的读取路径),并在服务边界注入十二种故障模式,包括延迟提交、重新投递和部分批次;每一轮实验都根据已提交效果的账本进行评分。在跨越九个近期模型、三个生产级智能体执行框架、两种契约变体和十五种恢复条件的25,930轮实验中,答案取决于故障类型。当立即回读能够揭示发生了什么时,模型决定结果:被指示精确执行一次的前沿模型几乎不会重复其确认丢失的写入(0.5%),较弱的模型则经常重复,模型解释了53%的已解释方差。当无法做到这一点时——请求仍在进行中,或传输层投递了两次——同样的前沿模型在56%和74%的轮次中发生重复,而契约解释了81%的方差。我们证明了在没有飞行时间上界的情况下,任何仅验证的策略在延迟提交下都无法实现精确一次。当这种上界较短且已知时,等待是有效的,但对于重尾的飞行中延迟,即使每轮等待一小时也不足以提供每次写入的幂等键,而提供幂等键可将重复率从28%降至4%,因为智能体在键存在时会使用它们。执行框架几乎无关紧要,一个附加键的守卫在不同执行框架间不变地转移,智能体在90%的已重复效果的轮次中报告了成功。
英文摘要
When a tool-using agent's write times out or returns a server error, the action may already have taken effect. Retrying blindly duplicates it -- a second charge, a second announcement, a second deployment -- while giving up skips required work. We ask where exactly-once behaviour should be enforced: in the model, in the agent harness, or in the tool contract. We introduce LIMBO, a deterministic sandbox of six services with realistic contracts (optional idempotency keys, eventually consistent and missing read paths) and twelve fault modes injected at the service boundary, including late commits, redelivery and partial batches; every episode is graded against a ledger of committed effects. Across 25,930 episodes spanning nine recent models, three production agent harnesses, two contract variants and fifteen recovery conditions, the answer depends on the fault. When an immediate read-back can reveal what happened, the model decides: frontier models instructed to act exactly once almost never duplicate a write whose acknowledgement was lost (0.5%), weaker models often do, and the model explains 53% of the explained variance. When it cannot -- the request is still in flight, or the transport delivered it twice -- the same frontier models duplicate in 56% and 74% of episodes, and the contract explains 81%. We prove that no verification-only policy is exactly-once under late commits without a bound on in-flight time. Waiting works when such a bound is short and known, but with heavy-tailed in-flight delays even an hour of waiting per episode falls short of offering an idempotency key on every write, which lowers the duplicate rate from 28% to 4% because agents use keys when they exist. The harness barely matters, a guard that attaches keys transfers across harnesses unchanged, and agents reported success in 90% of the episodes in which they had duplicated an effect.
Comments23 pages, 6 figures, 13 tables