arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越并行盲性:块式生成中的信息下限与模型差距

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

Xinwei Qiang, Xiang Fang, Chang Chen, Zaifeng Pan, Yue Guan, Yufei Ding

arXiv 2608.27339首次发表:更新:

发表机构

University of California San Diego(加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对块式生成的拒绝率,提出信息下限与模型差距的分离方法,在多领域目标上揭示了模型性能瓶颈,明确了短距离条件作用与生成质量的差异。

AI 中文摘要

块式生成(Block drafter)在一次前向传播中生成多个token,早于目标token被确定,其拒绝率由两类损失混合导致:块内路径信息缺失与可观测信息建模不完善。接受长度无法区分这两类损失,我们通过信息下限(指定条件顺序下的最小期望拒绝率)将二者分离,超出该下限的拒绝率即为模型差距。在四个领域、四个开放权重目标及一个前沿API目标的目标rollout中估计这两个指标,得到三个发现:第一,Qwen3-4B最终位置的全并行信息下限达0.286,即使最优生成器的单位置接受率也仅71%;第二,一个已确定的token可消除该下限的86%至100%,该局部性也通过独立的互信息分析得到验证;第三,当前生成器的拒绝率远高于其信息下限:最终位置的模型差距占DFlash拒绝率的43%至64%,占DSpark经oracle条件调整后拒绝率的85%至92%。这些发现明确了短距离条件作用的价值与生成质量的差异。

英文摘要

Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Measurements across four domains and four open-weight targets, plus floor estimates on a frontier API target, yield three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. Guided by this separation of short-range conditioning from proposal quality, we replace a context-independent predecessor correction with a prefix-attention head that reads committed context conditional on the predecessor. With supporting backbone components, the resulting Qwen3-4B drafter improves mean serving accepted length by $3.93\%$ over the released DSpark checkpoint across nine tasks, without increasing the conditioning order. Code and data: https://github.com/tie-pilot-qxw/specfloor; drafter checkpoints: https://huggingface.co/TIE-Pilot/dspark-attnconv-block7-qwen3-4b.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑