arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25961cs.CVcs.AI

用于视频级矛盾心理和犹豫识别的交互流知识引导多模态推理

Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition

Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy

首次发表
浏览论文内容

中文总结 AI 辅助

研究视频级矛盾心理和犹豫识别难题,提出PRISM-AH框架,通过冻结编码器、轻量级流模型处理多模态冲突,经窗口级注释监督、知识引导大语言模型推理,在测试中取得较好宏观F1,验证推理增益可转移。

中文摘要 AI 辅助

矛盾心理和犹豫(A/H)是健康行为改变延迟或放弃之前的冲突情感状态。视频级别的A/H识别困难,因为信号源于面部、声音、语言和身体模态之间及内部的不一致,且因人而异。提出的PRISM-AH框架将A/H视为随时间展开的多模态冲突。冻结的视觉、音频和文本编码器在短时间窗口内对齐并传入轻量级流模型,该模型对跨模态不和谐进行评分,预测下一个窗口以揭示犹豫惊喜信号,发现行为原型,并以参与者元数据为条件。密集的窗口级注释作为辅助目标监督模型,并针对宏观F1校准决策阈值。知识引导的大语言模型随后使用数据集的专家线索分类法对结构化证据进行推理,只有当验证性能提高时才在后期融合其判断。在525个视频的标记公共测试分区上,PRISM-AH的宏观F1为0.6133,而报告的零样本基线为0.2827。推理增益被验证可从验证转移到更大的测试分区。

英文摘要

Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change. Recognition of A/H at the video level is difficult, since the signal arises from disagreement across and within facial, vocal, linguistic, and bodily modalities, and manifests differently across individuals. The proposed PRISM-AH (Predictive Reasoning over Interacting Streams for Multimodal Ambivalence/Hesitancy Recognition), is a framework that treats A/H as a multimodal conflict that unfolds over time. Frozen vision, audio, and text encoders are aligned into short time windows and passed to a lightweight streaming model that scores cross-modal dissonance, predicts each next window to expose a hesitation surprise signal, discovers behaviour prototypes, and is conditioned on participant metadata. Dense window-level annotations supervise the model as an auxiliary objective, and the decision threshold is calibrated for macro F1. A knowledge-guided large language model then reasons over structured evidence using the expert cue taxonomy of the dataset, and its verdict is fused late only when validation performance improves. On the labelled public test partition of 525 videos, PRISM-AH attains a macro F1 of 0.6133, compared to the reported zero-shot baseline of 0.2827. The reasoning gain is validated to transfer from validation to the larger test partition.

补充信息

↑