arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Candor-LR:用于音视频语音识别的双人对话数据集

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte

arXiv 2609.10394首次发表:更新:

发表机构

Trinity College Dublin(都柏林圣三一学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有AVSR基准缺乏自然对话复杂性的问题,提出基于CANDOR语料库的双人对话基准Candor-LR,包含1,656个视频会议,并验证了视觉信息在对话场景中的补偿作用及跨域鲁棒性提升。

AI 中文摘要

当前的音视频语音识别(AVSR)基准,如LRS3,严重依赖于清晰、脚本化和排练过的语音。它们未能反映自然对话的复杂性,这种复杂性涉及重叠语音、自发轮流发言、非脚本化词汇和多变的声学条件。为了将该领域推向现实对话,我们引入了Candor-LR,一个源自CANDOR语料库的对话基准,该语料库包含1,656个自然双人视频会议。我们定制的数据准备流程分别产生了713.5小时、10.1小时和60.1小时的训练、验证和测试数据。在Candor-LR上评估预训练的AVSR模型显示,与LRS3相比,仅音频的准确率急剧下降,但视觉线索有效补偿,在Candor-LR上比在LRS3上带来了更大的性能提升。此外,在该语料库上训练显著提高了在干净和嘈杂条件下的跨域鲁棒性,因为其逼真的对话数据捕获了更广泛的音视频特征。我们开源了我们的流程以确保可复现性,将Candor-LR确立为对话式AVSR的一个具有挑战性的基准。

英文摘要

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, spontaneous turn-taking, unscripted vocabulary and variable acoustic conditions. To shift the field toward realistic dialogue, we introduce Candor-LR, a conversational benchmark derived from the CANDOR corpus of 1,656 natural dyadic videoconferences. Our custom data preparation pipeline yields 713.5, 10.1, and 60.1 hours of training, validation, and test data, respectively. Evaluating pretrained AVSR models on Candor-LR reveals that audio-only accuracy drops sharply compared to LRS3, but visual cues compensate effectively, driving much larger performance gains on Candor-LR than on LRS3. Furthermore, training on this corpus significantly improves cross-domain robustness under both clean and noisy conditions, as its realistic conversational data captures broader audio-video features. We open-source our pipeline to ensure reproducibility, establishing Candor-LR as a challenging benchmark for conversational AVSR.

CommentsAccepted to IEEE SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑