发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SchemaFill通过槽位并行推测解码,在生成工具调用时并行生成候选参数并由目标模型验证,显著提升端到端吞吐量,最高达4.05倍。
AI 中文摘要
LLM智能体通过生成结构化工具调用来与外部系统交互。给定用户请求、对话上下文和工具模式目录,工具调用模型必须选择工具并生成其参数,可能在单个响应中产生多个调用。标准自回归解码逐token生成这些调用,对于涉及多个调用或许多参数字段的请求会产生大量延迟。显式的参数结构提供了并行生成的机会,但后续的参数值可能依赖于前面的字段和调用,因此独立生成的值可能与目标模型的输出不同。我们提出了SchemaFill,一种通过槽位并行推测解码实现高效LLM工具调用的框架。SchemaFill并发生成未来的槽位值作为候选,无需预先知道实际的调用序列或参数值。跨越多个字段和调用的候选被拼接起来,由目标模型在实际输出前缀下进行验证。只有验证过的token被提交,当候选不一致时,目标模型提供修正。这应用了目标验证,同时利用了跨槽位和调用的并行性。在Glaive和BFCL上,SchemaFill相比自回归解码实现了高达4.05倍的端到端吞吐量提升。代码可在该https URL获取。
英文摘要
LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding generates these calls token by token, incurring substantial latency for requests involving multiple calls or many argument fields. The explicit argument structure offers opportunities for parallel generation, but later argument values may depend on preceding fields and calls, so independently generated values can differ from the target model's output. We present SchemaFill, a framework for efficient LLM tool calling through slot-parallel speculative decoding. SchemaFill generates future slot values concurrently as candidates, without requiring advance knowledge of the actual call sequence or argument values. Candidates spanning multiple fields and calls are concatenated for verification by the target model under the actual output prefix. Only verified tokens are committed, and the target supplies corrections when candidates disagree. This applies target verification while exploiting parallelism across slots and calls. On Glaive and BFCL, SchemaFill achieves up to a 4.05$\times$ improvement in end-to-end throughput over autoregressive decoding. Code is available at https://github.com/Czzzk/SchemaFill.