发表机构
Huawei Technologies Co., Ltd.; University of Science and Technology of China(华为技术有限公司; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AsymSpec是面向智能体大语言模型的非对称投机解码框架,通过轻量级草稿读全输入、大型验证模型用压缩视图,实现近90%完整上下文准确率,获1.3-1.7倍吞吐量加速且计算成本仅0.2-0.3倍。
AI 中文摘要
智能体大语言模型(Agentic LLM)流水线在检索、工具使用和多轮交互过程中,随着上下文的积累,推理成本不断攀升。为控制延迟,部署时通常会压缩输入,但这会降低任务准确率。投机解码(Speculative Decoding,SD)可在无损失的情况下加速生成,但它假设草稿模型(drafter)和验证模型(verifier)共享完全相同的上下文,这使得投机解码无法解决准确率与开销之间的权衡问题。我们提出AsymSpec,这是一种打破该对称性的非对称投机解码框架:轻量级草稿模型读取完整输入,而大型验证模型则基于压缩后的视图运行。草稿模型通过对数几率的对比δ融合(contrastive δ-fusion)引导验证模型,该过程由感知发散的接受门(divergence-aware acceptance gate)调节,以保持验证稳定性和高草稿接受率。在四项智能体能力和两项端到端智能体基准上进行评估,AsymSpec平均达到完整上下文准确率的约90%,在孤立文本能力上实现1.3至1.7倍的吞吐量加速,同时仅需0.2至0.3倍的计算成本。这些结果表明,当压缩丢弃关键推理信号时,非对称上下文访问可带来显著增益。
英文摘要
Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy. Speculative decoding (SD) accelerates generation losslessly, yet it assumes the drafter and verifier share an identical context, preventing SD from resolving the accuracy-overhead trade-off. We propose AsymSpec, an asymmetric speculative decoding framework that breaks this symmetry: a lightweight drafter reads the full input while the large verifier operates on the compressed view. The drafter steers the verifier via a contrastive $δ$-fusion of logits, modulated by a divergence-aware acceptance gate that preserves verification stability and high draft acceptance rates. Evaluated across four agentic capabilities and two end-to-end agent benchmarks, AsymSpec reaches $\approx 90\%$ of full-context accuracy on average, delivering $1.3$--$1.7\times$ throughput speedups at $0.2$--$0.3\times$ the compute cost on isolated text capabilities. These results show that asymmetric context access yields substantial gains precisely when compression discards critical reasoning signals.
CommentsEMNLP Main Conference 2026