FuseAlign:野外环境下的强制对齐
FuseAlign: Forced Alignment in the Wild
浏览论文内容
中文总结 AI 辅助
针对现有强制对齐评估低估实际难度的缺陷,提出改进指标、真实ASR评分协议及基准AlignBench,并引入基于Transformer的FuseAlign模型,通过联合上下文化和在线标签校正实现鲁棒对齐,显著优于基线。
中文摘要 AI 辅助
词级强制对齐用于估计每个转录词在音频录音中出现的时间位置。它是基于文本的媒体编辑、字幕生成、语音数据整理和语音学分析的基础。现有评估因依赖短时、清晰的语音、完美的转录文本以及会掩盖重大对齐错误的指标,而低估了强制对齐的难度。相比之下,现实世界的媒体和数据处理流程处理的是长时且多样化的录音。此外,强制对齐器通常处理的是易出错的自动语音识别(ASR)输出。我们通过改进的评估指标、针对真实ASR转录本的评分协议,以及涵盖多样化说话人、声学和文本条件的基准数据集AlignBench,来解决这些差距。我们进一步引入了FuseAlign,一种基于Transformer的对齐器,在大型伪标签语音上训练,并采用在线标签校正。FuseAlign对音频和文本进行联合上下文建模,以定位粗粒度词。随后,模型以毫秒级分辨率细化边界,并在无需基于词典或维特比解码的情况下,检测音频中缺失的转录词。在AlignBench上,FuseAlign显著优于所有基线,并在真实ASR转录本下保持稳健。消融实验表明,卷积上采样和EMA快照标签校正比模型属性(如参数数量)更为重要。
英文摘要
Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect transcripts, and metrics that obscure consequential alignment errors. In contrast, real-world media and data-processing pipelines operate on long and diverse recordings. Additionally, forced aligners often operate on error-prone automatic speech recognition (ASR) output. We address these gaps with improved evaluation metrics, a scoring protocol for real ASR transcripts, and AlignBench, a benchmark spanning diverse speaker, acoustic, and text conditions. We further introduce FuseAlign, a transformer-based aligner trained on large-scale pseudo-labeled speech with online label correction. FuseAlign performs joint contextualization of audio and text for the localization of coarse words. The model then refines boundaries at millisecond resolution and detects missing transcript words in the audio without lexicon-based or Viterbi decoding. On AlignBench, FuseAlign substantially outperforms all baselines and remains robust under real ASR transcripts. Ablations show that convolutional upsampling and EMA-snapshot label correction matter more than model properties such as parameter count.
发表机构
- Descript
机构由 AI 辅助整理,请以论文原文为准。