arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00814cs.CL

OoO-Spec:用于快速工具调用的乱序语义推测

OoO-Spec: Out-of-Order Semantic Speculation for Fast Tool Calling

Zhiheng Zhang, Mujie Xu, Feiyu Sun, Zhixin Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出OoO-Spec,采用Qwen3-0.6B辅助模型乱序预测工具调用的函数选择与参数,无需目标模型特定训练,在多基准测试中实现显著加速,性能优于ToolSpec及其他学习辅助模型。

中文摘要 AI 辅助

大型语言模型(LLMs)逐token生成工具调用,尽管通常可根据请求和工具架构并行预测函数选择与参数值。ToolSpec通过生成架构token并检索早期调用降低该成本,但无法提出前两者中不存在的请求特定值。我们提出OoO-Spec,该方法乱序计算这些缺失的语义。请求到达时,Qwen3-0.6B辅助模型在一次并行请求级波中预测函数选择和所有架构定义的参数槽,同时目标模型开始ToolSpec解码。运行时组合参数值,将生成的调用渲染为文本,并将其暴露给后续候选构建轮次。目标模型轮询时不阻塞,用自身分词器重新分词就绪的提示,且始终是唯一的验证和提交权威。该辅助模型通过LoRA在Qwen2.5-32B教师轨迹上训练一次,可在Qwen2.5、Qwen3和Llama目标模型上不变使用,无需针对目标模型的特定辅助模型训练。在7个完全排名的目标模型和3个基准测试的贪心批量单解码设置下,OoO-Spec在全部21个目标-基准单元中是所有评估方法中最快的,与自回归解码相比达到2.46倍至5.34倍的加速,未加权平均加速为3.89倍,而ToolSpec的加速为2.95倍。它在每个可比单元中也优于所有评估的已发布学习辅助模型。在Qwen3-4B、8B、14B和32B目标模型上,同一辅助模型较ToolSpec平均提升34.1%。其紧凑的语义有效载荷平均每个请求85字节(不含协议元数据),支持有效的GPU拆分重叠。

英文摘要

LLMs generate tool calls token by token, even though the function choice and argument values can often be predicted in parallel from the request and tool schema. ToolSpec reduces this cost by drafting schema tokens and retrieving earlier calls, but cannot propose request-specific values absent from either source. We present OoO-Spec, which computes these missing semantics out of order. At request arrival, a Qwen3-0.6B sidecar predicts the function choice and all schema-defined argument slots in one parallel request-level wave while the target begins ToolSpec decoding. The runtime joins the slot values, renders the resulting call as text, and exposes it to subsequent candidate-construction rounds. The target polls without blocking, re-tokenizes a ready hint with its own tokenizer, and remains the sole verifier and commit authority. The sidecar is trained once with LoRA on Qwen2.5-32B teacher traces and used unchanged across Qwen2.5, Qwen3, and Llama targets, without target-specific drafter training. Across seven fully ranked targets and three benchmarks under greedy batch-one decoding, OoO-Spec is fastest among all evaluated methods in all 21 target-benchmark cells, reaching 2.46x-5.34x over autoregressive decoding with an unweighted mean of 3.89x, versus 2.95x for ToolSpec. It also outperforms every evaluated released learned drafter in each comparable cell. Across Qwen3-4B, 8B, 14B, and 32B targets, the same sidecar improves on ToolSpec by 34.1% on average. Its compact semantic payload averages 85 bytes per request excluding protocol metadata, supporting effective split-GPU overlap.

补充信息

↑