arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21827cs.CLcs.LG

RheoSampling:解决随机动态树投机解码中的独热困境

RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding

Qiao Hu, Yepeng Weng, Bo Zhang, Takehisa Yairi

首次发表
浏览论文内容

中文总结 AI 辅助

针对动态树投机解码在随机采样时因独热分布导致接受率下降的问题,提出RheoSampling方法,通过解耦树构建与令牌验证的角色,实现上下文感知与随机采样的无损结合,提升推理效率。

中文摘要 AI 辅助

投机解码通过并行草拟多个令牌来加速大语言模型推理,基于树的方法通过层次结构进一步提高效率。诸如EAGLE-3之类的动态树方法在贪婪解码下通过确定性top-K扩展和全局剪枝表现良好。然而,在随机解码(T>0)中,该机制将草稿分布压缩为独热概率,导致接受率严重下降。这造成了一个困境:动态树方法牺牲随机采样以保留上下文感知的拓扑结构,而静态树方法以上下文无关的结构保留随机采样。问题源于同一概率分布被用于两个冲突的任务:构建树和验证令牌。这种耦合使得直接注入随机性因所产生的随机过程而具有挑战性。我们通过解耦这些角色来解决此问题:RheoSampling为从草稿分布中采样的令牌分配一个代理概率用于树扩展和剪枝,同时保留其真实采样概率用于验证。具体而言,我们在确定性top-K槽中注入一个采样令牌,并在构建和验证期间以不同概率处理它,使RheoSampling成为首个同时具备上下文感知top-K构建和随机采样且保持无损性的动态树方法。我们通过等价类分析建立了无损保证,该分析将随机树空间压缩为可处理的类别。基于最优传输的验证策略和稀疏草稿机制确保理论收益转化为实际效率。跨大语言模型和基准的实验表明,在接受率和加速比上优于最先进的动态树方法。该框架可能为分析随机树结构提供模板。

英文摘要

Speculative decoding accelerates LLM inference by drafting multiple tokens in parallel, with tree-based methods further improving efficiency through hierarchical structures. Dynamic-tree methods such as EAGLE-3 perform well under greedy decoding via deterministic top-K expansion and global pruning. However, in stochastic decoding (T>0), this mechanism collapses the draft distribution into one-hot probabilities, causing a severe drop in acceptance rate. This creates a dilemma: dynamic-tree methods sacrifice stochastic sampling to preserve context-aware topology, while static-tree methods preserve stochastic sampling with context-agnostic structures. The issue arises because the same probability distribution is used for two conflicting tasks: constructing the tree and verifying tokens. This coupling makes direct injection of randomness challenging due to the resulting stochastic process. We resolve this by decoupling these roles: RheoSampling assigns a token sampled from the draft distribution a proxy probability for tree expansion and pruning alongside its true sampling probability for verification. Specifically, we inject a sampled token among the deterministic top-K slots and treat it with different probabilities during construction and verification, making RheoSampling the first dynamic-tree method with both context-aware top-K construction and stochastic sampling while maintaining losslessness. We establish the lossless guarantee through an equivalence-class analysis that compresses the stochastic tree space into tractable classes. An OT-based verification strategy and a sparse draft mechanism ensure that theoretical gains translate into practical efficiency. Experiments across LLMs and benchmarks demonstrate improvements in acceptance rate and speedup over state-of-the-art dynamic tree methods. This framework may provide a template for analyzing stochastic tree structures.

发表机构

  • National Center for Mathematics and Interdisciplinary Sciences (NCMIS), AMSS, CAS(中国科学院数学与系统科学研究院国家数学与交叉科学中心)
  • The University of Tokyo(东京大学)
  • SKLMS and AMSS, Chinese Academy of Sciences(中国科学院数学与系统科学研究院,科学与工程计算国家重点实验室)
  • School of Mathematical Sciences, University of Chinese Academy of Sciences(中国科学院大学数学科学学院)
  • Lenovo AI Technology Center(联想AI技术中心)

机构由 AI 辅助整理,请以论文原文为准。

↑