发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LiLiCorr是关联并行草稿边际分布的轻量级模型,联合训练草稿生成器后,在多数设置下提升了推测解码的接受长度与吞吐量,且对长输入仍保持性能优势。
AI 中文摘要
推测解码通过生成未来token并由目标模型并行验证,加速语言模型推理。DFlash这类扩散式块头是极具吸引力的草稿生成器,它在一次前向传播中预测一整段未来token。然而,DFlash是基于每个位置的边际分布而非联合块分布进行训练的,因此其生成的token单独来看合理,但整体上不连贯。我们提出LiLiCorr,这是一种轻量级似然模型,用于关联草稿生成器已产生的每个位置的边际分布。它保留每个位置的前k个token作为候选,对其进行联合处理,为每个候选生成入向量和出向量。当较早候选的出向量与较晚候选的入向量具有高余弦相似度时,一对相邻候选即匹配。这些匹配捕获了块的联合结构,而无需显式构建完整的联合分布。一次轻量级网络前向传播即可生成所有向量,随后成对分数以批量矩阵运算并行计算,仅留下廉价的贪心搜索作为顺序操作。我们进一步将草稿生成器与LiLiCorr联合训练,使其学会提出能关联成更长接受序列的候选。在普通DFlash草稿生成器上,LiLiCorr将所有基准测试的接受长度提高了9%至19%,而其评分头仅占每个块延迟的约2.8%。与DFlash及另外两种在草稿阶段恢复连贯性的同期方法相比,在72种设置中的70种设置下,LiLiCorr实现了最高吞吐量:包括贪心解码和温度为1的解码下的9个基准测试、两种目标大小,以及在6种并发度、两种输入长度和3种熵层级上的吞吐量扫描,所有系统均在通用服务栈上进行了同等优化。将LiLiCorr扩展至比其训练时长一个数量级更长的输入时,仍保持了这一领先优势。
英文摘要
Speculative decoding accelerates language-model inference by drafting future tokens the target model verifies in parallel. A diffusion-style drafter such as DFlash drafts an entire block in one forward pass. It is trained on the per-position marginals rather than on the joint distribution over the block, so the tokens it emits are individually plausible yet jointly incoherent. We introduce LiLiCorr, a Lightweight Likelihood-based model that Correlates the per-position marginals such a drafter produces. It keeps the top-K tokens at each position and processes them jointly, emitting an in and an out vector for each. Two candidates at consecutive positions match when the earlier out vector aligns, in cosine similarity, with the later in vector. Training scores the correct pairings highest and pushes competing ones down, so coherent blocks outscore incoherent ones. The joint distribution over the block, exponential in its length, is never materialized. One lightweight network pass produces all the vectors, the pairwise scores follow as batched matrix operations, leaving only a cheap greedy walk sequential. We co-train the DFlash drafter with LiLiCorr, so it proposes candidates that correlate into longer accepted sequences. Over the vanilla DFlash drafter it builds on, LiLiCorr accepts more and serves faster at all 72 settings we test: nine benchmarks at two target sizes under greedy and temperature-one decoding, plus a throughput sweep over six concurrencies, two input lengths and three output-entropy tiers. It raises acceptance length by 7 to 19%, while its single-pass scoring head costs only about 3% of the per-block latency. Against three concurrently developed methods that also restore coherence at draft time, all equally optimized on a common stack, LiLiCorr holds the highest throughput in 63 of those settings, ties within a measured noise floor in 6, and trails in only 3.