REDDIT:基于回放的分布编辑在无遗忘条件下校正ASR的模型生成时间戳漂移
REDDIT: Forgetting-Resistant Correction of Timestamp Drift in ASR via Replay-Based Distribution Editing
- National Taiwan University(台湾大学)
- Carnegie Mellon University(卡内基梅隆大学)
- NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(国立清华大学人工智能研究中心(NTU AI-CoRE))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对自回归ASR的非语音区间时间戳漂移问题,提出REDDIT两阶段后训练框架,仅用少量数据和参数更新校正时间戳,避免普通微调的灾难性遗忘。
AI中文摘要:
现代自回归自动语音识别(ASR)系统可将时间戳作为解码令牌输出,无需帧级对齐器或推理时后处理即可生成带时间戳的转录结果。本文发现这类生成的时间戳会在长非语音区间发生漂移:转录内容可能仍合理,但解码出的时间轴会与音频实际时间偏离。我们基于自建的间隙与长间隙基准测试集,在15个可生成时间戳的ASR及音频语言系统上研究了这种非语音引发的时间戳漂移问题。朴素的时间戳校正微调可提升对齐效果,但会严重损害非目标ASR的性能,存在遗忘问题。本文提出REDDIT(基于回放的分布编辑),一种轻量级两阶段后训练框架,可在校正时间戳的同时避免灾难性遗忘:首先在模型自身回放的解码器上下文下编辑时间戳目标,同时在非时间戳令牌上匹配冻结的基础分布,随后执行简短的编辑前缀精调阶段。该框架结合经语音活动检测(VAD)裁剪的语音区间、插入的非语音间隙与已知拼接偏移量,无需人工转录或人工时间戳标注即可构建校正监督信号。在Whisper-tiny模型上,仅使用34.9小时的目标校正音频、更新1.6%的模型参数,即可将长间隙交并比均值(mIoU)从38.7%提升至95.0%,将混合间隙域外平均绝对偏移量(AAS)从2752毫秒降至223毫秒,同时将CV-en的词错误率(MER)维持在41.3%,而普通监督微调(SFT)解码器调优的该指标高达524.2%。
英文摘要:
Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or inference-time post-processing. We show that these generated timestamps can drift across long non-speech spans: the transcript may remain plausible, but the decoded time axis drifts away from the audio. We study this non-speech-induced timestamp drift with self-built gap and long-gap benchmarks across 15 evaluated timestamp-producing ASR and audio-language systems. Naive timestamp-corrected fine-tuning improves alignment but can severely degrade non-target ASR behavior, exposing a forgetting problem. We propose REDDIT(REplay-based Distribution eDITing), a lightweight two-stage post-training framework that corrects timestamps while avoiding this catastrophic forgetting: it first edits timestamp targets under the model's own replayed decoder context while matching the frozen base distribution on non-timestamp tokens, then applies a short edited-prefix refinement stage. In this framework, we construct correction supervision without human transcripts or human timestamp annotations by combining VAD-trimmed speech spans with inserted non-speech gaps and known concatenation offsets. On Whisper-tiny, 34.9 hours of targeted correction audio used and only 1.6% of model parameters updated, raising long-gap mIoU from 38.7% to 95.0% and reducing mixed-gap out-of-domain AAS from 2752 ms to 223 ms while preserving CV-en MER at 41.3% (versus 524.2% for ordinary SFT decoder tuning).