AI 中文总结
该研究针对ASR单模型推测解码的对齐漂移问题,分析其机制,提出两种修正方法,其中AnchorDraft可提升端到端速度,揭示ASR自推测依赖的关键因素。
AI 中文摘要
推测解码通过让轻量草稿模型提出多个token,由目标模型一次检查来加快生成速度。在单模型形式中,草稿是附加在目标模型上的轻量模块,而非独立模型。将该设计应用于自动语音识别(ASR)会引入额外问题:草稿每一步都可读取完整音频,但单独运行时其提案会变差,访问并非定位。已接受文本明确记录转录位置,而草稿还需跟踪不断变化的音频位置。在主要匹配对比中,每步音频访问对首个提案的改变不大,但会使后续提案的接受率大致翻倍。固定宽度窗口显示,音频位置可解释部分差距:位置正确的窗口能恢复连续性,而宽度相同但位置错误的窗口会降低连续性。在报告的最严苛条件下,草稿后期的中位误差达21帧,而验证阶段的目标注意力始终保持在2帧中位范围内。我们测试了两种减少漂移的方法:第一种从验证注意力中读取音频位置,用于指导下一轮草稿,仅当额外接受的token抵消读出成本时才能节省时间;第二种是AnchorDraft,它在训练时教草稿跟踪音频位置,无需修改推理图,训练后的草稿在两种测试的目标规模下均提升了端到端速度。这些结果表明,ASR自推测依赖于token预测、音频位置跟踪和草稿成本。
英文摘要
Speculative decoding speeds up generation by letting a cheap draft propose several tokens that a target model checks in one pass. In the single-model form, the draft is a lightweight module attached to the target rather than a separate model. Applying this design to Automatic Speech Recognition (ASR) introduces an extra problem. The draft can read the whole audio at every step, yet its proposals get worse as it runs on its own. Access is not localization. The accepted text keeps the transcript position explicit, but the draft must also track the changing audio position. In the primary matched comparison, per-step audio access changes the first proposal modestly but roughly doubles later-proposal acceptance. Fixed-width windows show that the audio position explains part of this gap. A correctly placed window recovers continuation, while an equally narrow window at the wrong position reduces it. Late-draft median error reaches 21 frames in the hardest reported condition, while target attention during verification stays within a 2-frame median. We test two ways to reduce this drift. The first reads the audio position from verification attention and uses it to guide the next draft round. It saves time only when the extra accepted tokens offset the readout cost. The second is AnchorDraft, which teaches the draft to track the audio position during training without changing the inference graph. The trained draft improves end-to-end speed at both tested target scales. These results show that ASR self-speculation depends on token prediction, audio-position tracking, and draft cost.