$C^3$ASD:多层次一致性驱动的表示学习
$C^3$ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection
浏览论文内容
中文总结 AI 辅助
研究视频中活跃说话者检测问题,提出$C^3$ASD框架,含嵌入级、序列级和预测级一致性约束,提升跨模态表示鲁棒性,在多种损坏下有显著改进。
中文摘要 AI 辅助
活跃说话者检测确定视频中可见人物在每个时刻是否在说话。近期视听融合方法在干净数据上表现良好,但在现实世界损坏(如背景噪声、遮挡或同时的模态退化)下会退化。我们将此限制归因于缺乏促进跨模态稳健、语义对齐表示的显式一致性约束。我们提出了$C^3$ASD,一个具有三个互补约束的多层次一致性驱动框架。广泛的实验表明,在各种音频、视觉和联合损坏下有显著改进,同时在干净数据上保持有竞争力的性能。
英文摘要
Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrade under real-world corruptions such as background noise, occlusion, or simultaneous modality degradation. We attribute this limitation to the absence of explicit consistency constraints that promote robust, semantically aligned representations across modalities. Without such guidance, models tend to learn fragile modality-specific shortcuts that fail under corrupted conditions. We propose $C^3$ASD, a multi-level consistency-driven framework with three complementary constraints: embedding-level inter-modality consistency aligns audio-visual representations during speech; sequence-level intra-modality consistency separates speaking and non-speaking clusters via track-aware contrastive learning; and prediction-level consistency stabilizes fusion through knowledge distillation. Extensive experiments demonstrate significant improvements under diverse audio, visual and joint corruptions, while maintaining competitive performance on clean data.
发表机构
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。