arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

声学进度传播用于ASR长程投机解码

Acoustic Progress Propagation for Long-Horizon Speculative Decoding in ASR

Yuanyuan Jia, Qianqian Yang

arXiv 2609.33245首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出进度感知投机草稿模型,通过循环传播声学进度状态指导令牌生成,在五个ASR测试集上实现无损加速,优于基线。

AI 中文摘要

投机解码可加速自回归自动语音识别(ASR),但对齐感知草稿模型的接受长度会随着草稿视界增加而饱和。我们提出一种进度感知的投机草稿模型,该模型在草稿步骤间循环传播声学进度状态,并将其反馈到音频交叉注意力中以指导令牌生成。我们在可变草稿视界上联合训练草稿模型和进度预测器。在五个ASR测试集上,我们的方法相对于仅目标自回归解码,在Qwen3-ASR-0.6B和Qwen3-ASR-1.7B上分别实现了无损的宏平均端到端加速1.657倍和1.227倍。与AnchorDraft相比,我们的方法将宏平均加速分别提高了34.3%和9.0%。视界扫描显示,接受长度在基线饱和点之后持续增长。代码可在该HTTPS URL获取。

英文摘要

Speculative decoding accelerates autoregressive automatic speech recognition (ASR), but the acceptance length of alignment-aware drafters can saturate as the draft horizon increases. We propose a progress-aware speculative drafter that recurrently propagates an acoustic progress state across draft steps and feeds it back into audio cross-attention to guide token generation. We jointly train the drafter and progress predictor over variable draft horizons. On five ASR test sets, our method achieves lossless, macro-averaged end-to-end speedups of 1.657x and 1.227x over target-only autoregressive decoding with Qwen3-ASR-0.6B and Qwen3-ASR-1.7B, respectively. Relative to AnchorDraft, our method improves the macro-averaged speedup by 34.3% and 9.0%, respectively. Horizon sweeps show continued growth in acceptance length beyond the baselines' saturation. Code is available at https://github.com/yuanyuanjia71-spec/ProgDraft.

Comments5 pages, 2 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑