发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgSpec通过补充语料库和自适应草稿长度策略,提升了编码智能体流程中基于检索的推测解码性能,在仓库级多智能体基准上最高实现4.76倍吞吐量提升。
AI 中文摘要
基于检索的推测解码(SD)通过从现有文本中复制续写来草拟令牌,这适用于编码智能体,因为它们会反复生成代码、日志和之前的尝试。然而,现有方法在智能体流程中存在不足:大量可复用文本要么不在其语料库中,要么存储形式与智能体输出的形式不同,而且它们的草稿长度忽略了接受长度在不同智能体间存在差异并随轮次漂移的情况。我们提出了AgSpec,一个为现有检索引擎在编码智能体流程中所缺乏的语料库和草稿长度策略提供支持的框架。AgSpec从会话、工作区和全局语料库中检索,保留正在进行的会话轨迹,并以智能体的输出格式对打开的文件进行索引。它通过离线分析的上限来限制每个智能体的草稿长度,并根据验证反馈在线调整长度。在两个仓库级多智能体编码基准上,AgSpec在大多数评估设置中优于五种基于检索的草稿生成器和EAGLE-3,在批大小为1时将生成吞吐量相对于自回归解码提升至4.37倍,在批大小为16时提升至4.76倍。AgSpec在没有仓库或多智能体流程的基准上仍然有效,表明其增益广泛适用于编码智能体。
英文摘要
Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37$\times$ at batch size 1 and 4.76$\times$ at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.