发表机构
nyra labs(nyra实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现代语音识别模型中因转录风格为潜在变量导致的问题,通过覆盖感知解码器任务令牌等方法,提升德语不流畅F1,全英文微调超越基线,引入监督交叉注意力微调改进时间戳,还提出新任务用于创建和丰富语音语料库。
AI 中文摘要
现代在异构注释数据上训练的语音识别模型将转录风格(逐字与意图)视为不可控的潜在变量,导致可测量的解码不稳定、评估混淆(高达60%的词错误率归因于风格不匹配)和不可靠的词级定时。研究表明模型已对两种风格进行编码,挑战在于可控激活。使用在平行逐字/意图转录对上训练的覆盖感知解码器任务令牌,德语不流畅F1零样本从10%提高到79%。全英文微调在逐字准确性、不流畅检测和两种语言的意图模式质量上超越所有基线。还引入监督交叉注意力微调改进不流畅语音的词级时间戳。最后提出verbatimize新任务,可扩展创建和丰富具有高质量规范逐字转录的语音语料库。
英文摘要
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.
CommentsAccepted at Interspeech 2026 long track