arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00881cs.LG

AOSpec:用于低延迟智能体服务的动作与观测协同推测

AOSpec: Action and Observation Co-Speculation for Low-Latency Agent Serving

Hao Mark Chen, Jinnan Guo, Wayne Luk, Hongxiang Fan

首次发表
浏览论文内容

中文总结 AI 辅助

AOSpec是一种动作与观测协同推测的无损框架,通过EVD和JASV技术降低智能体服务延迟,在Terminal-Bench上较基线实现显著延迟缩减,且观测模型可跨基准迁移。

中文摘要 AI 辅助

大型语言模型智能体越来越多地通过有状态工具执行操作,但模型生成与环境执行在每一步都保持串行化。随着解码速度提升,工具执行正成为日益突出的瓶颈。现有仅针对动作或观测的推测方法会暴露大量延迟:价值集中在少数慢速调用中,部分结果仅能通过执行产生,而更长的前瞻通常需要越来越不可能的动作预测链。我们提出AOSpec,这是一种在完整智能体-环境循环中协同推测动作与观测的无损框架。期望价值解码(EVD)将观测推测导向具有最大预期延迟收益的结果,优化隐藏的预期时间而非命中率。对于仅能通过执行揭示的结果,AOSpec在包含其效果的隔离分支中启动延迟关键的目标动作,同时联合动作-状态验证(JASV)在重用前针对已提交执行验证动作及其源状态。JASV将长程动作依赖从全链预测重构为目标动作-状态验证,打破前瞻-准确率权衡并解锁长程重叠,且不牺牲串行语义。在覆盖四个测试框架、五个动作模型和五个服务速度的Terminal-Bench服务设置中,AOSpec优于所有实用基线,将平均端到端延迟降低11.8-32.5%,p99延迟最多降低42.8%。其增益随解码加速而增加,且其观测模型无需重新训练即可从Terminal-Bench迁移至SWE-bench Verified。

英文摘要

Large language model agents increasingly act through stateful tools, yet model generation and environment execution remain serialized at every step. As decoding accelerates, tool execution becomes a growing bottleneck. Existing action- or observation-only speculation leaves much of this latency exposed: value is concentrated in a few slow calls, some outcomes emerge only through execution, and longer lookahead typically requires an increasingly unlikely chain of action predictions. We present AOSpec, a lossless framework that co-speculates actions and observations across the full agent-environment loop. Expected Value Decoding (EVD) directs observation speculation toward outcomes with the greatest expected latency benefit, optimizing expected time hidden rather than hit rate. For outcomes only execution can reveal, AOSpec launches latency-critical target actions in isolated forks that contain their effects, while Joint Action-State Verification (JASV) verifies both the action and its origin state against committed execution before reuse. JASV recasts long-horizon action dependency from full-chain prediction into target action-state verification, breaking the lookahead--accuracy tradeoff and unlocking long-range overlap without sacrificing serial semantics. Across Terminal-Bench serving settings spanning four harnesses, five actor models, and five serving speeds, AOSpec outperforms every practical baseline, reducing mean end-to-end latency by 11.8-32.5% and p99 latency by up to 42.8%. Its gains increase as decoding accelerates, and its observation model transfers from Terminal-Bench to SWE-bench Verified without retraining.

↑