arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31170cs.CL

面向WhisperX的上下文感知交错批处理

Context-Aware Interleaved Batching for WhisperX

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Carlos Bain, Max Bain

AI总结:

针对WhisperX音频内批处理丢失上下文、标准Whisper推理慢且易幻觉的问题,提出上下文感知交错批处理,利用VAD边界稳定Whisper文本条件,在长音频基准上降WER、提升专有名词转录且保持高推理速度。

AI中文摘要:

WhisperX通过音频内批处理加速语音转录,但它会隔离音频片段,丢失连贯标点和术语转录所需的历史上下文;相反,标准Whisper按顺序保留上下文,但存在推理速度慢和幻觉循环的问题。为融合两者优势,我们提出上下文感知交错批处理,该算法利用语音活动检测(VAD)生成的片段边界,稳定Whisper的文本条件,使我们能在批处理音频片段间安全维持连续历史上下文。在长音频基准测试中,该方法降低了词错误率(WER),提升了专有名词转录效果,同时保持了高吞吐量推理速度。

英文摘要:

While WhisperX accelerates speech transcription via intra-audio batching, it isolates audio segments, losing the historical context needed for coherent punctuation and terminology transcription. Conversely, standard Whisper retains context sequentially but suffers from slow inference and hallucination loops. To achieve the best of both worlds, we propose Context-Aware Interleaved Batching. By using VAD-derived segment boundaries, our algorithm stabilizes Whisper's text conditioning, allowing us to safely maintain continuous historical context across batched audio segments. As demonstrated on long-form audio benchmarks, this approach reduces Word Error Rate (WER) and improves proper noun transcription, all while maintaining high-throughput inference speeds.

↑