arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaSpark: 具有在线学习的自适应DSpark用于树验证与N-gram填充

AdaSpark: Adaptive DSpark with Online Learning for Tree Verification and N-gram Fill

Liquan Liu, Yifan Zhang, Bowei Xu

arXiv 2610.05774首次发表:更新:

发表机构

Zeraix(Zeraix)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AdaSpark是一种在线学习调度器,无需预配置即可自适应选择验证宽度和候选排序,通过定价时间与接受概率,在多种目标上实现1.5-3.1倍加速,且不损失精度。

AI 中文摘要

诸如DSpark之类的块草稿器在一次前向传播中为多个位置提出排名候选,而树验证器在目标的一次传递中检查它们。要验证的行数权衡了更宽的树预期接受的令牌数与更宽的验证所需时间。大多数选择此数量的调度器从服务前测量的表或模型中获取验证时间,并通过最多一个缩放因子在线校正,并从草稿器的置信度估计或离线拟合的映射中获取接受率。AdaSpark在服务过程中学习这两个量,无需预先进行配置、校准或扫描。它学习哪些验证宽度值得提供,并将每个宽度的验证时间拟合为上下文的函数。它将每个候选的接受概率拟合到目标的验证结果,以草稿器的置信度头作为输入之一,并根据该拟合而不是置信度头来排序和调整树的大小。同一模型对请求自身文本的n-gram延续进行定价,因此草稿和文本衍生的候选在一个最佳优先顺序中竞争行数。宽度通过以长期解码速率定价时间来选择。在来自六个公共数据集的单轮和多轮对话中,针对三个密集目标和一个人工专家混合目标,AdaSpark的解码速度比this http URL的DSpark使用相同草稿器快1.5-3.1倍。我们的imparo引擎与AdaSpark比imparo运行三令牌链(默认的this http URL设置)快1.17-1.52倍;这一增益仅来自调度器。在没有宽度扫描的情况下,AdaSpark在任何密集目标或上下文带上从未比最佳固定树宽度慢超过0.3%。在人工专家混合目标上,它与最佳固定宽度持平,而其他从4到16行的固定宽度则慢5-14%。

英文摘要

Block drafters such as DSpark propose ranked candidates for several positions in one forward pass, and a tree verifier checks them in one pass of the target. The number of rows to verify trades the tokens a wider tree is expected to accept against the time a wider verify takes. Most schedulers that choose this number take the verify time from a table or model measured before serving, corrected online by at most one scale factor, and take acceptance from the drafter's confidence estimates or from a map fitted offline. AdaSpark learns both quantities while it serves, with no profile, calibration or sweep in advance. It learns which verify widths are worth offering and fits each one's verify time as a function of context. It fits each candidate's acceptance probability to the target's verify outcomes, with the drafter's confidence head as one input, and orders and sizes the tree by that fit instead of by the head. The same model prices n-gram continuations of the request's own text, so drafted and text-derived candidates compete for rows in one best-first order. The width is chosen by pricing time at the long-run decode rate. On single- and multi-turn conversations from six public datasets, on three dense targets and one mixture-of-experts target, AdaSpark decodes 1.5-3.1x faster than llama.cpp's DSpark with the same drafters. Our imparo engine with AdaSpark is 1.17-1.52x faster than imparo running with a three-token chain (the default llama.cpp setting); this gain comes from the scheduler alone. Without a width sweep, AdaSpark is never more than 0.3% slower than the best pinned tree width on any dense target or context band. On the mixture-of-experts target it ties the best pinned width, and the other pinned widths from 4 to 16 rows are 5-14% slower.

Comments25 pages, 10 figures, 15 tables. Code: https://github.com/zeraix/imparo

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑