发表机构
Tencent Inc.(腾讯公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在通过推测解码实现高性能推理,提出AngelSpec统一训练框架,在训练、架构、推理层面解决异质性问题,其DFly框架提升了接受长度和吞吐量。
AI 中文摘要
推测解码可加速大语言模型推理且不改变目标分布,但没有单一的起草结构能在所有现实工作负载中表现最佳。自回归多令牌预测(MTP)是一种轻量级、稳定的提议机制,而块并行扩散可在更长的候选序列上分摊起草延迟。我们提出AngelSpec,这是一个统一的训练框架,在三个层面解决这种异质性。在训练层面,共同专门化结构和数据;在架构层面,提出DFly框架;在推理层面,将验证视为共享的批处理级资源。实验表明,DFly在Hy3系列上提高了平均接受长度,提升了吞吐量。我们发布AngelSpec以支持这些方法的训练和扩展。
英文摘要
Speculative decoding accelerates large language model inference without changing the target distribution, but no single drafting structure performs best across real-world workloads. Autoregressive multi-token prediction (MTP) is a lightweight, stable proposal mechanism, whereas block-parallel diffusion amortizes drafting latency over much longer candidate sequences; the better choice depends strongly on the output distribution. We present AngelSpec, a unified training framework for MTP and block-parallel speculative decoding that addresses this heterogeneity at three levels. At the training level, rather than fitting one universal drafter to a uniform data mixture, we co-specialize structure and data: the MTP drafter is trained on diverse conversational data for high-entropy open-ended chat, and the block-diffusion drafter on code and mathematics data for longer predictable continuations. At the architecture level, we propose DFly, a block-diffusion framework combining a hybrid target-conditioning backbone with a predecessor-conditioned autoregressive head, improving target-feature utilization and intra-block dependency modeling while keeping generation parallel. At the inference level, both acceptance length and verification cost vary with domain, request, online load, and hardware, so DFly treats verification as a shared batch-level resource: it reallocates compute toward high-confidence prefixes across requests and combines expected utility with a profiled cost model to adapt verification depth online. Across the Hy3 series, DFly raises the average accepted length on Hy3-A21B by roughly 30% and attains the highest average throughput at every tested concurrency from 4 to 64, a 1.98-2.40x speedup over autoregressive decoding and 10.5-11.8% higher throughput than DFlash. We release AngelSpec to support training and extending these methods.