arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

复制还是不复制:通过内在模型信号控制推测解码

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik, Lior Wolf, Itamar Zimerman

arXiv 2609.20186首次发表:更新:

发表机构

Tel Aviv University; Stealth Startup, Tel Aviv(特拉维夫大学; 特拉维夫隐秘初创公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对推测解码中神经草稿与复制策略的权衡,提出SwitchSD框架,利用轻量级探针识别LLM内部复制意图,动态切换策略,在Llama和Qwen上实现高达15%的吞吐量提升。

AI 中文摘要

推测解码(Speculative Decoding, SD)显著加速了大语言模型(LLM)的推理过程,然而现有方法在两种草稿生成策略之间面临根本性权衡:神经草稿生成与基于上下文的复制。神经草稿(如EAGLE3)在多样化的文本设置中提供稳健的性能,而基于复制的方法通过在复制密集型场景中更快地生成候选词并利用长重复片段实现近乎完美的推测,从而获得更高的加速比。我们分析了现有的基于复制的方法,发现它们容易产生偶然重复,其中表面上的n-gram重叠并不反映结构性的复制意图,导致误触发,最终降低吞吐量。我们引入了SwitchSD,一个自适应框架,将复制视为LLM的潜在控制信号。通过在目标模型的内部表示上训练轻量级探针,SwitchSD以高精度(AUC > 0.99)识别真实的复制意图。这使得系统能够在神经草稿(如EAGLE)和基于上下文的复制之间动态切换。我们在Llama和Qwen系列上的结果表明,与最先进的基线(如EAGLE3)相比,吞吐量提升高达15%,有效地将复制从一种噪声启发式方法转变为一种有原则的、模型感知的解码机制。

英文摘要

Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text settings, while copy-based methods achieve higher speedups in copy-intensive regimes by generating candidates faster and exploiting long repetition spans for near-perfect speculation. We analyze existing copy-based methods and find that they are prone to accidental repetitions where surface-level n-gram overlap does not reflect a structural intent to copy, leading to false-positive triggers that ultimately degrade throughput. We introduce SwitchSD, an adaptive framework that treats copying as a latent control signal of the LLM. By training lightweight probes on the target model's internal representations, SwitchSD identifies genuine copy-intent with high precision (AUC > 0.99). This allows the system to dynamically switch between neural drafting (e.g., EAGLE) and context-based copying. Our results across Llama and Qwen families demonstrate throughput gains of up to 15% over state-of-the-art baselines like EAGLE3, effectively turning copying from a noisy heuristic into a principled, model-aware decoding regime.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑