AI 中文总结
Oilbird提出利用验证器已计算的隐藏状态重新键控上下文池的语义草稿源,提升无需训练的推测解码的接受长度与速度,在API-Bank上达4.4倍自回归解码速度。
AI 中文摘要
无需训练的推测解码通过将上下文的精确后缀与早期上下文池匹配来生成草稿。这种查找会遗漏池中已有的正确草稿,在工具调用场景中表现尤为明显:请求几乎重复所有内容,仅生成少数新值,而一个被拒绝的标记会丢弃其背后的正确延续。我们在十个基准测试中逐位置诊断失败,发现问题在于寻址而非覆盖:在最密集的工具调用基准测试中,最强的精确匹配草稿生成器遗漏的内容中,约一半存在于池中但无法通过精确匹配获取。因此,我们提出第二种语义草稿源:同一上下文池,通过验证器在每个已提交标记处已计算的隐藏状态重新键控,并结合合并机制使其可嵌入现有词汇草稿生成器的树结构中。在三个已发布的草稿生成器中,在匹配的池和预算下,该方法将接受的生成长度提升了24%-29%。在API-Bank上,Oilbird实现了4.4倍的自回归解码速度,而我们测试中最强的无需训练基线为3.9倍,EAGLE-3为2.0倍。
英文摘要
Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct drafts already in the pool, most visibly on tool-calling traffic, where a request repeats almost everything but the few values minted for it, and where one rejected token discards the correct continuation behind it. We diagnose the failure position by position across ten benchmarks and find it to be a problem of addressing rather than of coverage: on our densest tool-calling benchmark, about half of what the strongest exact-match drafter misses is present in the pool yet unreachable by exact matching. We therefore propose a second, semantic draft source: the same pool, re-keyed by the hidden state the verifier has already computed at each committed token, together with a merge that lets it ride inside an existing lexical drafter's tree. In three published drafters, at matched pool and budget, it lifts accepted length by 24-29%. Oilbird reaches 4.4x autoregressive decoding speed on API-Bank, against 3.9x for the strongest training-free baseline in our harness and 2.0x for EAGLE-3.