arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24004cs.CL

AgentSpec:面向大语言模型智能体批量推理的推测解码算法

AgentSpec: Speculative Decoding for Batch Inference of LLM Agents

Xin Wang, Ziming Miao, Yi Zhu, Hui Shen, Zhongwei Wan, Fan Yang, Mi Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体批量推理响应慢的问题,本研究提出AgentSpec推测解码算法,通过结构隔离的draft机制和感知冗余的预算分配提升效率,在多工作负载与模型上验证其性能优于现有最优方法。

中文摘要 AI 辅助

基于大语言模型(LLM)的智能体应用往往存在响应时间长的问题,推测解码是一种可在不影响生成质量的前提下提升LLM智能体推理效率的有前景方案。然而,现有最优的推测解码算法在大批量规模下会出现显著的速度下降,限制了其在实际智能体应用中的部署效果。本研究首先对LLM智能体的推测解码进行系统分析,识别出加速比下降的两个主导因素:推测令牌的高拒绝率,以及动态令牌预算的利用不足。基于这些发现,我们提出AgentSpec算法,这是一种针对LLM智能体的推测解码算法,用于解决现有方法的局限。AgentSpec包含结构隔离的 draft(推测)机制,将推测范围限制在智能体工作流的语义连贯片段中,减少无关语义路径的 draft,实现极低的拒绝率;还采用感知冗余的预算分配机制,利用智能体层级信息,在推理过程中更好地利用动态空闲令牌预算。我们在vLLM框架上针对5种不同工作负载、4个来自4个不同LLM家族的模型实现并评估AgentSpec,结果表明AgentSpec优于现有最优方法。

英文摘要

Large language model (LLM)-based agent applications often incur high response time. Speculative decoding is a promising solution to improve the inference efficiency of LLM agents without impacting generation quality. However, state-of-the-art speculative decoding algorithms exhibit substantial speed degradation under large batch sizes, limiting their effectiveness to deploy in real-world agent applications. In this work, we first present a systematic analysis of speculative decoding for LLM agents and identify two dominant factors of speedup degradation: high rejection rate of speculative tokens, and under-utilization of dynamic token budgets.B ased on these observations, we propose AgentSpec, a speculative decoding algorithm that addresses the limitations of existing methods for LLM agents. AgentSpec incorporates structure-isolated drafting that constrains speculation to semantically coherent segments of the agent workflow, reducing the drafts of irrelevant semantic paths and achieving an extremely low rejection rate. Moreover, AgentSpec adopts redundancy-aware budget allocation that exploits agent-level information to better utilize the dynamically-free token budget during the agent inference. We implement and evaluate AgentSpec on five different workloads and four different models from four different LLM families in vLLM. Our results demonstrate the superiority of AgentSpec over state-of-the-arts.

发表机构

  • The Ohio State University(俄亥俄州立大学)
  • Microsoft Research(微软研究院)
  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑