arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21347cs.CV

质量感知多模态融合揭示效价-唤醒特征中的隐含身份

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

Jisu Kim, Benjamin S. Riggan

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对传统人脸识别在无约束环境中表现不佳的问题,提出质量感知自适应融合方法 QAAF 用于多模态效价-唤醒估计,在 VA 估计任务中取得良好效果,且经 VA 训练的特征在编码身份上表现出色,确立了多模态 VA 估计为传统人脸识别的互补软生物特征模态。

中文摘要 AI 辅助

传统人脸识别依赖静态外观线索,在无约束环境中表现不佳。我们假设视听表达动态携带与静态外观互补的身份判别信息,提取此信号需要对野外视频可变输入质量具有鲁棒性的多模态表示。为此,我们将多模态效价-唤醒(VA)估计作为一个 pretext 任务,并提出质量感知自适应融合(QAAF),通过学习软门控和质量依赖的随机失活来估计每个样本、每个模态的可靠性并调整每个模态的贡献。对于 VA 估计问题,QAAF 在 Aff-wild2 上通过后期融合集成实现了平均一致性相关系数(CCC)为 0.472,优于相同设置下的基线集成(0.415)和单骨干基线(0.288)。此外,QAAF 对不可用模态具有更强的弹性,当一个模态缺失时,CCC 相对下降仅 7.5 - 34.4%。然后我们探究这些经过 VA 训练的特征是否无需特定身份训练就能编码身份。在 AFEW-VA(67 个演员)和 YTF(1595 个受试者)上,经过 VA 训练的骨干特征在评估的软生物特征方法中排名第一,并且与 ArcFace 的分数级融合降低了两个数据集上的错误接受率(AFEW-VA 上从 0.022 降至 0.021,YTF 上从 0.106 降至 0.104),纠正了 AFEW-VA 上 ArcFace 的 68.2%的错误接受。这些发现确立了多模态 VA 估计作为与传统人脸识别互补的软生物特征模态。

英文摘要

Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lighting. We hypothesize that audiovisual expression dynamics carry identity-discriminative information complementary to static appearance, and that extracting this signal requires multimodal representations robust to the variable input quality of in-the-wild video. To learn such representations, we cast multimodal valence-arousal (VA) estimation as a pretext task and propose Quality-Aware Adaptive Fusion (QAAF), which estimates per-sample, per-modality reliability and adapts each modality's contribution through learned soft gating and a quality-dependent dropout. For the problem of VA estimation, QAAF achieves an average Concordance Correlation Coefficient (CCC) of 0.472 via late fusion ensembling on Aff-wild2, improving over a baseline ensemble under the same setting (0.415) as well as a single-backbone baseline (0.288). Furthermore, the proposed QAAF demonstrates greater resilience to unavailable modalities, with only a 7.5-34.4% relative decrease in CCC when one modality is missing. We then probe whether these VA-trained features encode identity without identity-specific training. On AFEW-VA (67 actors) and YTF (1,595 subjects), VA-trained backbone features rank first among evaluated soft biometric methods, and score-level fusion with ArcFace lowers EER on both datasets (0.022 to 0.021 on AFEW-VA, 0.106 to 0.104 on YTF), correcting 68.2% of ArcFace's false accepts on AFEW-VA. These findings establish multimodal VA estimation as a soft biometric modality complementary to conventional face recognition.

发表机构

  • University of Nebraska-Lincoln(内布拉斯加大学林肯分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑