FFASR:使用高保真模拟RIR的远场自动语音识别基准测试
FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs
浏览论文内容
中文总结 AI 辅助
FFASR通过高保真模拟RIR构建远场ASR基准,覆盖多种声学条件,验证模拟可替代实测评估。
中文摘要 AI 辅助
远场自动语音识别(ASR)在混响、噪声和说话人运动条件下性能下降,然而驱动模型选择的基准测试却强调近讲麦克风语音。我们提出FFASR,一个包含15,637条话语的留出语料库和一个覆盖九种条件的开放排行榜,每种条件仅改变一个声学因素:无回声近场语音、实测与模拟办公室实验室配对、高/中/低信噪比(SNR)下的静态远场混合语音,以及在匹配SNR下的移动说话人变体。来自15位说话人的干语音与来自14个配备家具房间的混合波动/几何声学房间冲激响应进行卷积;由于语音是新录制的且测试波形从未发布,该语料库能抵抗训练数据污染。在当代系统中,平均词错误率(WER)从近场的4.4%上升到静态低SNR条件下的41.3%;在匹配SNR下,移动说话人带来虽小但一致的惩罚;在办公室实验室配对中,实测和模拟WER平均相差约1.7个百分点。这些结果支持高保真模拟作为我们所测试条件下实测远场评估的可扩展替代方案。
英文摘要
Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying a single acoustic factor: anechoic near-field speech, a measured-versus-simulated office-lab pair, static far-field mixtures at high/mid/low signal-to-noise ratio(SNR), and moving-talker variants at matched SNR. Dry speech from 15 talkers is convolved with hybrid wave/geometrical-acoustics room impulse responses from 14 furnished rooms; because the speech is newly recorded and the test waveforms are never released, the corpus resists training-data contamination. Across contemporary systems, mean word error rate (WER) rises from 4.4% near-field to 41.3% in the static low-SNR condition; a moving talker adds a small but consistent penalty at matched SNR; and on the office-lab pair, measured and simulated WER agree to within about 1.7 pp on average. These results support high-fidelity simulation as a scalable proxy for measured far-field evaluation under the conditions we test.
发表机构
- Hugging Face
- Treble Technologies
机构由 AI 辅助整理,请以论文原文为准。