arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38201cs.CLcs.OScs.SE

TomasuLLM:面向LLM智能体的乱序推测执行

TomasuLLM: Out-of-Order Speculative Execution for LLM Agents

  • HKUST(香港科技大学)
  • Nanjing University(南京大学)
  • King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
  • Zhejiang Lab(之江实验室)

机构由 AI 辅助整理,请以论文原文为准。

Jiangnan Yu, Ceyu Xu, Mengming Li, Shiyu Huang, Yiran Xia, Jian Weng, Hui Xue, Haohui Mai, Yuan Xie

AI总结:

TomasuLLM通过乱序推测执行和写时复制沙箱验证,在保持正确性的前提下,将编码智能体的工具调用延迟转化为加速,在多个基准上实现1.27至1.35倍的性能提升。

AI中文摘要:

长时间运行的工具可能主导编码智能体的延迟:编译器、测试套件和仓库命令需要数秒到数分钟的时间,而智能体却处于空闲状态。这种观察停滞现象呈现出与推动乱序处理器发展的相同张力——顺序接口隐藏了可以预测并提前启动的工作,但推测结果可能只有在它本身及所有更早步骤都经过验证之后才变得可见。我们提出TomasuLLM,一个在保持任务执行正确性的同时,按轨迹顺序之外执行智能体工具调用的运行时。它草拟未来的动作,在隔离的写时复制沙箱中运行它们,追踪它们的依赖关系和效果,并仅在针对已提交状态进行验证后,按轨迹顺序提交结果。在三个跨越亚秒级到分钟级工具调用的基准测试中,TomasuLLM提升了报告的基准平均值,并随工具延迟扩展:在100个SWE-bench Verified任务上提升1.31倍,在28个Terminal-Bench 2.0任务上提升1.35倍,在18个SWE-Marathon会话上实现1.27倍的匹配进度。在4,010条审计的提交验证记录中,它产生了零个误接受。

英文摘要:

Long-running tools can dominate coding-agent latency: compilers, test suites, and repository commands take seconds to minutes while the agent idles. This observation stall presents the same tension that drove out-of-order processors -- asequential interface hides work that can be predicted and started early, but a speculative result may become visible only after it and every earlier step have been validated. We present TomasuLLM, a runtime that executes agent tool calls out of trajectory order while preserving task-execution correctness. It drafts future actions, runs them in isolated copy-on-write sandboxes, traces their dependencies and effects, and commits results in trajectory order only after validation against committed state. Across three benchmarks spanning sub-second to minutes-long tool calls, TomasuLLM improves the reported benchmark means and scales with tool latency: 1.31x on 100 SWE-bench Verified tasks, 1.35x on 28 Terminal-Bench 2.0 tasks, and 1.27x matched progress on 18 SWE-Marathon sessions. Across 4,010 audited commit-validation records, it produces zero false accepts.

↑