arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AngelSpec:通过推测解码实现现实世界中的高性能推理

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

Hong Liu, Rui Cen, Junhan Shi, Guangshuo Qin, Jiebin Zhang, Tianyu Liu, Runzhi Fan, Guoliang Zhao, Ruobing Xie, Kai Zhang, Song Liu, Guanghua Yu, Jianchen Zhu

arXiv 2607.25852首次发表:更新:

发表机构

Tencent Inc.(腾讯公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在通过推测解码实现高性能推理,提出AngelSpec统一训练框架,在训练、架构、推理层面解决异质性问题,其DFly框架提升了接受长度和吞吐量。

AI 中文摘要

推测解码可加速大语言模型推理且不改变目标分布,但没有单一的起草结构能在所有现实工作负载中表现最佳。自回归多令牌预测(MTP)是一种轻量级、稳定的提议机制,而块并行扩散可在更长的候选序列上分摊起草延迟。我们提出AngelSpec,这是一个统一的训练框架,在三个层面解决这种异质性。在训练层面,共同专门化结构和数据;在架构层面,提出DFly框架;在推理层面,将验证视为共享的批处理级资源。实验表明,DFly在Hy3系列上提高了平均接受长度,提升了吞吐量。我们发布AngelSpec以支持这些方法的训练和扩展。

英文摘要

Speculative decoding accelerates large language model inference without changing the target distribution, but no single drafting structure performs best across real-world workloads. Autoregressive multi-token prediction (MTP) is a lightweight, stable proposal mechanism, whereas block-parallel diffusion amortizes drafting latency over much longer candidate sequences; the better choice depends strongly on the output distribution. We present AngelSpec, a unified training framework for MTP and block-parallel speculative decoding that addresses this heterogeneity at three levels. At the training level, rather than fitting one universal drafter to a uniform data mixture, we co-specialize structure and data: the MTP drafter is trained on diverse conversational data for high-entropy open-ended chat, and the block-diffusion drafter on code and mathematics data for longer predictable continuations. At the architecture level, we propose DFly, a block-diffusion framework combining a hybrid target-conditioning backbone with a predecessor-conditioned autoregressive head, improving target-feature utilization and intra-block dependency modeling while keeping generation parallel. At the inference level, both acceptance length and verification cost vary with domain, request, online load, and hardware, so DFly treats verification as a shared batch-level resource: it reallocates compute toward high-confidence prefixes across requests and combines expected utility with a profiled cost model to adapt verification depth online. Across the Hy3 series, DFly raises the average accepted length on Hy3-A21B by roughly 30% and attains the highest average throughput at every tested concurrency from 4 to 64, a 1.98-2.40x speedup over autoregressive decoding and 10.5-11.8% higher throughput than DFlash. We release AngelSpec to support training and extending these methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑