arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SVF-CR:用于多模态矛盾心理和犹豫识别的同步视觉-面部交叉细化

SVF-CR: Synchronized Visual-Facial Cross-Refinement for Multimodal Ambivalence and Hesitancy Recognition

Hyein Park, Namho Kim, Junhwa Kim

arXiv 2607.09417首次发表:更新:

发表机构

Konyang University; Korean Broadcasting System (KBS)(公州大学; 韩国广播公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态矛盾心理和犹豫识别,提出SVF-CR框架,通过提取视频和面部令牌,经多种注意力机制细化,构建视觉-面部证据并融合文本、声学特征,实验证明该框架提升了公共宏F1。

AI 中文摘要

矛盾心理和犹豫是通过言语内容、面部行为、视觉背景和声学线索组合表达的微妙行为状态。有效识别不仅需要提取信息丰富的单模态表示,还需对跨模态时间对齐的行为证据交互进行建模。本文提出了一种具有成对多模态证据融合的同步视觉-面部交叉细化框架(SVF-CR)用于矛盾心理和犹豫识别。该方法首先使用相同时间划分提取全视频段令牌和裁剪面部段令牌,通过模态内自注意力和双向视觉-面部交叉注意力对同步的视觉和面部令牌进行细化。然后利用一致性和差异建模构建段级视觉-面部证据,接着进行时间自注意力和注意力池化。文本和声学特征通过上下文自注意力进行轻度细化,并在最终决策阶段使用成对证据融合与增强的视觉-面部证据融合。在BAH公共评估分割上的实验表明,所提出的同步视觉-面部交叉细化在全局视觉-面部令牌融合和同步证据基线之上提高了公共宏F1,达到了0.7156的公共宏F1。代码可在指定链接获取。

英文摘要

Ambivalence and hesitancy are subtle behavioral states that are expressed through a combination of verbal content, facial behavior, visual context, and acoustic cues. Effective recognition therefore requires not only extracting informative unimodal representations, but also modeling how temporally aligned behavioral evidence interacts across modalities. In this paper, we propose a synchronized visual-facial cross-refinement framework (SVF-CR) with pairwise multimodal evidence fusion for ambivalence and hesitancy recognition. The proposed method first extracts whole-video segment tokens and cropped-face segment tokens using the same temporal partition. The synchronized visual and facial tokens are refined through intra-modal self-attention and bidirectional visual-facial cross-attention, allowing whole-video context and local facial behavior to mutually refine each other before evidence construction. We then construct segment-level visual-facial evidence using consistency and discrepancy modeling, followed by temporal self-attention and attention pooling. Textual and acoustic features are lightly refined through context self-attention and are fused with the enhanced visual-facial evidence at the final decision stage using pairwise evidence fusion. Experiments on the BAH (Behavioral Ambivalence/Hesitancy) public evaluation split show that the proposed synchronized visual-facial cross-refinement improves public macro-F1 over both global visual-face token fusion and synchronized evidence baselines, achieving a public macro-F1 of 0.7156. Code is available at : https://github.com/hiinnnii/BAH-Challenge-ECCV2026\_SVF-CR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑