发表机构
Trinity College Dublin(都柏林圣三一学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出AVSRBench多条件基准,评估三种AVSR架构在六种条件下的表现,揭示视觉性能退化及泛化差距,并引入RoomReader-AV基准及统一预处理流水线。
AI 中文摘要
虽然音视频语音识别(AVSR)在标准LRS3基准上已实现低于1%的词错误率,但其对广播语音的依赖掩盖了这是否反映真正的泛化能力还是仅属于领域适应。为探究这一差距,我们评估了三种AVSR架构在六种条件下的表现:受控广播语音、固定语法话语、过度清晰的伦巴第语音、来自专业唇语者及非专业演讲者的朗读语音,以及自发的多方视频对话。我们发现,纯视觉性能在广播领域之外迅速恶化,而音视频融合主要惠及伦巴第语音环境。在90°侧面视角下,视觉理解急剧下降,多模态系统主要依赖声学回退。此外,说话者的发音清晰度比轻微相机位移更为关键,且基于大语言模型(LLM)的架构在域外泛化方面表现不佳。我们的工作凸显了当前AVSR研究中显著的泛化差距。为解决此问题,我们还引入了RoomReader-AV作为新的AVSR基准,并发布了一个统一的数据预处理流水线,使全面的多条件评估变得可行。
英文摘要
While AVSR has achieved sub-1% word error rates on the standard LRS3 benchmark, its reliance on broadcast speech obscures whether this reflects true generalization or just domain adaptation. To investigate this gap, we evaluate three AVSR architectures across six conditions: controlled broadcast speech, fixed-grammar utterances, hyper-articulated Lombard speech, read speech from professional lipspeakers and non-professional speakers, and spontaneous multi-party video conversations. We find that visual-only performance deteriorates rapidly beyond broadcast domains, and audio-video fusion mainly benefits Lombard speech environments. Visual understanding degrades sharply at 90° profile views, with multimodal systems relying largely on acoustic fallback. Additionally, speaker articulation proves more critical than minor camera shifts, and LLM-based architectures suffer from poor out-of-domain generalization. Our work highlights a significant generalization gap in current AVSR research. To address this, we also introduce RoomReader-AV as a new benchmark for AVSR and release a unified data preprocessing pipeline to make comprehensive multi-condition evaluation accessible.
CommentsAccepted to IEEE SLT 2026