arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FA-Bench:在干净与噪声条件下进行词级和音素级强制对齐及ASR时间戳的基准测试

FA-Bench: A Benchmark for Phone- and Word-Level Timestamp Accuracy in Forced Alignment and ASR on Clean and Noisy Speech

Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han, Yichao Zhou, Ruchao Fan, Bingshen Mu, Jingbei Li, Vishwas Shetty, Sarthak Bisht, Ziyue Qiu, Massa Baali, Rita Singh, Bhisha Raj

arXiv 2609.32396首次发表:更新:

AI 中文总结

FA-Bench是一个开放基准,统一协议下评估21个开放模型和9个商业API在干净和噪声条件下的强制对齐与ASR时间戳,采用容差F1消除分数膨胀,并发现Whisper提前150毫秒等系统性计时偏差。

AI 中文摘要

强制对齐将语音音频与文本转录对齐,以生成词和音素的时间戳。已发表的比较在归一化转录、数据划分和边界匹配方式上各不相同,因此其数字无法相互对照。我们提出了FA-Bench,一个开放框架,它固定了这些选择,并发布了代码、数据划分、音素映射、文本归一化和评分脚本,结果定期公布。轨道1为每个对齐器提供参考转录,轨道2提供识别器的输出,在相同的音频上,以四种方式降质,在统一协议下使用21个开放模型和9个商业API进行测试。我们对话语的每个边界进行评分,并检查其旁边的两个标签,因此识别器遗漏或虚构的词会被计分。使用基于容差的F1作为主要指标,消除了标准MAE在会话语音中对识别依赖系统造成的9%至14%的分数膨胀。然后我们按边界位置和相邻词被正确识别的数量对边界进行分组,这显示了系统在何处丢失了分数。我们发现了当前系统对词计时方式的系统性偏差,Whisper大约提前150毫秒,而几个商业ASR API延迟超过50毫秒。代码和结果位于此https URL。

英文摘要

Forced alignment estimates the timestamps of each word, phone or character in speech given its transcript. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices once and releases the code, splits, phone mapping, text normalization and scoring script, with results published periodically. Track 1 gives every aligner the reference transcript and Track 2 gives it a recognizer's output, on the same audio, clean and degraded four ways, with 21 open models and 9 commercial APIs under a unified protocol. We score every boundary of an utterance and check the two labels beside it, so a word the recognizer missed or invented is charged. Using a tolerance-based F1 as our primary metric eliminates the 9% to 14% score inflation that standard MAE causes on recognition-dependent systems in conversational speech. We then group boundaries by position and how many adjacent words were recognized correctly, which shows where a system lost the score. We discovered systematic bias in how current systems time words, with Whisper about 150 ms early and several commercial ASR APIs over 50 ms late. Code and results are at https://github.com/olewave/fa-bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑