arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Emo-DVS:面向隐私感知情感识别的事件相机多模态基准

Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras

Jiaqi Chen, Qinfu Xu, Hao Zhuang, Liyuan Pan

arXiv 2609.06928首次发表:更新:

发表机构

Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RGB相机隐私风险及现有事件方法数据有限、单模态局限,提出三模态基准Emo-DVS及信息引导门控融合(IGF)框架,通过预训练、自适应门控和互信息最大化实现最先进性能。

AI 中文摘要

情感分析是计算机视觉中的一项基本任务,但其实际部署一直受到传统RGB相机固有隐私风险的限制。受生物启发的事件相机提供了一种有前景的硬件级解决方案,因为它们捕获异步亮度变化,从而减少面部身份细节的暴露,同时利用高动态范围在具有挑战性的光照条件下实现稳健感知。尽管有这些优势,现有基于事件的方法在复杂的现实世界环境中仍面临困难,原因在于数据集规模有限、采集条件简单以及依赖单模态视觉线索。为解决这些问题,我们构建了一个具有挑战性的三模态基准,包含事件、音频和文本模态,并提出了信息引导门控融合(IGF)框架,该框架首先在Emo-DVS的FAU子集上预训练事件编码器以捕获细粒度的面部动态,然后采用自适应模态门控来抑制模态特定噪声,最后利用互信息最大化来对齐稳健的跨模态表示。为缓解数据稀缺,我们引入了Emo-DVS,这是首个大规模基于事件的情感分析数据集,它将动态光照与面部动作单元(FAU)子集和情感子集相结合。大量实验表明,IGF实现了最先进的性能。

英文摘要

Emotion analysis is a fundamental task in computer vision, but its practical deployment remains constrained by the privacy risks inherent to conventional RGB cameras. Bio-inspired event cameras present a promising hardware-level solution because they capture asynchronous brightness changes, thereby reducing exposure of facial identity details while leveraging high dynamic range for robust perception under challenging illumination conditions. Despite these advantages, existing event-based methods struggle in complex real-world settings due to limited dataset scales, simple acquisition conditions, and reliance on single-modality visual cues. To address these, we establish a challenging tri-modal benchmark with event, audio, and text modalities and propose the Information-Guided Gated Fusion (IGF) framework, which first pre-trains an event encoder on the FAU subset of Emo-DVS to capture fine-grained facial dynamics, then employs adaptive modality gating to suppress modality-specific noise, and finally leverages mutual information maximization to align robust cross-modal representations. To alleviate data scarcity, we introduce Emo-DVS, the first large-scale event-based emotion analysis dataset, which couples dynamic illumination with the Facial Action Unit (FAU) subset and emotion subset. Extensive experiments demonstrate that IGF achieves state-of-the-art performance.

Comments8 pages,5 figures,Accepted to ACM Multimedia 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑