arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

导师解码:更快的推理遇上提升

Mentored Decoding: Faster Inference meets Boosting

Vivien Tran-Thien, Richard Nock

arXiv 2609.30474首次发表:更新:

发表机构

Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出导师解码,一种有损推测解码的正式方法,通过连接推理与提升理论,证明其能在加速推理的同时提升质量,并推广至所有f-散度,实现高效的最优参数查询与分布构造。

AI 中文摘要

推测解码是一种成功的技术,通过快速的起草模型加速目标自回归语言模型的推理。有损推测解码允许与目标存在偏差,以进一步提高速度。有趣的是,实验观察到所得模型在质量上也能超越目标。我们的论文正式证明了这种壮举如何可能,通过一种称为导师解码的正式方法来实现有损推测解码。为此,我们将推理与著名的机器学习训练理论——提升——联系起来,并将导师解码推广到整个f-散度集合。我们揭示了导师解码的关键性质,包括(i)全变差情形特别吸引人的几何性质,(ii)与提升合规性直接相关的任何f-散度的简单近似,以及(iii)一种与散度无关的O(n)空间和O(sort(n))时间的数据结构,基于起草者和目标输出构建,允许在O(log n)时间内查询对偶问题的最优参数,并在O(n)时间内为任何f-散度构造最优导师分布。

英文摘要

Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can $\textit{also}$ beat the target $\textit{quality-wise}$. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called $\textit{mentored decoding}$. To get there, we connect inference to a celebrated ML training theory, $\textit{boosting}$, and proceed via the generalization of mentored decoding to the whole set of $f$-divergences. We uncover key properties of mentored decoding, among which (i) the particularly appealing geometric nature of the total variation case, (ii) simple approximations for any $f$-divergence in direct relation with boosting compliance, and (iii) a $\textit{divergence independent}$ $O(n)$ space and $O(\mathrm{sort}(n))$ time data structure built on drafter and target outputs, which allows to query the optimal parameters of the dual problem in $O(\log n)$ time and constructing optimal mentored distributions in $O(n)$ time for any $f$-divergence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑