arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02331cs.CVcs.AI

用于野外身体情绪表达的上下文感知领域专家混合模型

Context-Aware Mixture of Domain Experts for Bodily Expression of Emotion in the Wild

Mohammad Mahdi Dehshibi, David Masip

首次发表
浏览论文内容

中文总结 AI 辅助

提出CA-MoDE模型,将场景与物体线索作为结构化先验,通过最大认可门控策略融合多领域信号,在单张静态图像上实现优于现有时间模型的身体情绪识别性能。

中文摘要 AI 辅助

相同的身体姿态可根据周围上下文传达完全不同的情绪,但大多数身体情绪识别方法将场景和物体线索视为辅助特征增强,而非情绪合理性的结构化先验。我们提出用于身体情绪识别的上下文感知领域专家混合模型(Context-Aware Mixture of Domain Experts, CA-MoDE),该模型包含专门的场景专家与物体专家,可基于各自领域生成情绪类别的软分布,这些领域条件下的软预测作为结构化上下文先验,在分布层面而非特征层面调节身体专家的预测。为融合多领域信号,我们提出任务定制的最大认可门控策略,该策略为每个情绪维度选择专家中最强的上下文信号,缓解了对冲突或无信息上下文分布进行平均时通常出现的信号稀释问题。CA-MoDE在身体语言数据库上实现了0.3269的情绪识别得分,仅用单张静态图像就优于现有时间模型,表明显式建模结构化空间上下文可作为视频通常捕获的行为动态的补充判别代理。

英文摘要

The same body posture can convey entirely different emotions depending on its surrounding context, yet most methods for recognising bodily emotions treat scene and object cues as auxiliary feature augmentations rather than as structured priors over the plausibility of emotions. We introduce the Context-Aware Mixture of Domain Experts (CA-MoDE) for bodily emotion recognition. CA-MoDE incorporates dedicated scene and object experts to generate soft distributions over emotion categories conditioned on their respective domains. These domain-conditioned soft predictions serve as structured contextual priors that modulate the body expert's predictions at the distributional level rather than at the feature level. To fuse these multi-domain signals, we propose a task-tailored max-endorsement gating strategy that selects the strongest contextual signal across experts for each emotion dimension. Our gating strategy mitigates the signal dilution that typically occurs when conflicting or uninformative context distributions are averaged. CA-MoDE achieves an Emotion Recognition Score of 0.3269 on the Body Language Database. By outperforming existing temporal models using only single still images, our framework demonstrates that explicitly modelling structured spatial context can serve as a complementary discriminative proxy for the behavioural dynamics typically captured by video.

发表机构

  • University of the West of England(西英格兰大学)
  • Universitat Oberta de Catalunya(开放加泰罗尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑