arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于加快使用工具的智能体的推测性宏提交(Speculative Macro Commit)

Speculative Macro Commit for Faster Tool-Using Agents

Zeyu Liu, Souvik Kundu, Peter A. Beerel

arXiv 2609.03236首次发表:更新:

发表机构

University of Southern California; Intel Labs(南加州大学; 英特尔实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SMC机制,通过双层智能体系统复用多步推测性执行,在保持准确率的同时显著降低工具使用智能体的延迟,相关代码已公开。

AI 中文摘要

使用工具的大语言模型(LLM)智能体的实际耗时不仅消耗在模型推理上,还消耗在串行的动作-观测轮次中,每一次工具调用、环境转换和观测都会延迟后续决策。本文提出了推测性宏提交(Speculative Macro Commit,SMC),这是一种用于双层智能体系统的运行时机制:一个大型权威执行模型生成官方轨迹,而一个速度更快的推测性起草模型在隔离的环境快照上持续预测并执行未来的动作链。SMC从训练轨迹中挖掘重复出现的多动作框架,并将其存储在宏库中,用于在运行时匹配起草模型预测的动作链。当权威模型的下一次工具调用与第一个起草的动作匹配时,SMC会将剩余的预执行起草步骤及其观测结果提交到官方轨迹。以Qwen3.5-27B INT4作为权威执行模型,Qwen3.5-4B作为推测性起草模型,SMC在保持顺序智能体整体准确率的同时,在τ²-Bench电信子集上比推测性动作(Speculative Actions,SA)基线减少10.23%的延迟,比顺序执行减少18.59%;在AppWorld上,SMC比SA基线减少7.7%的实际耗时,比顺序执行减少44.9%,仅伴随任务完成率的小幅下降。总体而言,SMC提供了一种实用的方法来复用多步推测性执行,减少超出单步推测性动作的智能体延迟,其代码可在此处获取。

英文摘要

Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and observation can delay subsequent decisions. We introduce \textbf{Speculative Macro Commit} (SMC), a runtime mechanism for a two-tier agent system: a large authoritative actor model produces the official trajectory, while a faster speculative drafter model continuously predicts and executes future action chains on an isolated environment snapshot. SMC mines recurring multi-action skeletons from training traces and stores them in a macro library used to match against action chains predicted by the drafter at runtime. When the actor's next tool call matches the first drafted action, SMC commits the remaining pre-executed draft steps, together with their observations, to the official trajectory. Using Qwen3.5-27B INT4 as the authoritative actor model and Qwen3.5-4B as the speculative drafter model, SMC matches the sequential agent's overall accuracy while reducing latency by 10.23\% over the Speculative Actions (SA) baseline and 18.59\% over sequential execution on the $τ^2$-Bench Telecom subset. On AppWorld, SMC reduces wall time by 7.7\% over SA baseline and 44.9\% over sequential execution, with a small reduction in task completion. Overall, SMC provides a practical way to reuse multi-step speculative execution and reduce agent latency beyond single-step speculative actions. Our code is publicly available \href{https://github.com/zeyuliu1037/speculative-macro-commit}{\textcolor{magenta}{here}}.

CommentsAccepted in MLSP2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑