发表机构
KU Leuven; Xiaomi Corporation(鲁汶大学; 小米公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ROAM-ASD,一种联合建模音频、全脸和嘴部表示并采用统一自注意力与模态丢弃的视听框架,在五个ASD基准上取得最先进性能,并显著提升零样本泛化与鲁棒性。
AI 中文摘要
活动说话人检测(ASD)要求可见人脸与声学语音之间建立可靠的关联,然而现有系统在具有挑战性的领域或不完整的观测条件下常常性能下降。我们提出了ROAM-ASD,一个鲁棒的视听框架,它联合建模音频、全脸和细粒度嘴部表示。一个统一的联合自注意力机制将所有输入流与模态无关的查询令牌一起处理,使得可用的模态输入之间能够直接交互。模态丢弃进一步提高了输入流不可用时的鲁棒性。ROAM-ASD在五个ASD基准上达到了最先进的性能:在WASD上达到98.8%的mAP,在UniTalk上达到87.9%,在AVA上达到96.5%,在ASW上达到99.3%,在Talkies上达到98.2%,相比之前的最佳系统分别提高了5.1、4.7、0.9、1.0和2.1个mAP点。ROAM-ASD还大幅提升了零样本跨数据集泛化能力,并且对缺失观测保持鲁棒。
英文摘要
Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.
CommentsSubmitted to IEEE ICASSP 2027