发表机构
Harbin Institute of Technology; SERES(哈尔滨工业大学; 赛力斯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有车内情感计算数据集局限,引入多模态InCarEmo数据集,整合多种数据支持情感识别、疲劳检测和分心监测等任务,构建中英文基准并给出基线结果,证明多模态融合优势,为车内情感理解奠基,推动人车交互发展。
AI 中文摘要
理解驾驶员的情绪和状态对于确保安全并增强人车交互的下一代智能车内系统至关重要。然而,现有的车内情感计算公共数据集大多局限于视觉模态,很少包含对话信息,难以捕捉驾驶员情绪背后的语言和交互线索。为填补这些空白,我们引入了InCarEmo,一个用于车内情感识别和驾驶员状态监测的多模态数据集。它整合了RGB和红外视频、车内音频以及从模拟现实驾驶员行为的脚本化车内场景收集的对话文本,涵盖多种光照条件和驾驶环境。该数据集支持多模态情感识别、疲劳检测和分心监测这三项主要任务。除了原始中文数据,还构建了辅助英文基准以支持初步跨语言评估。我们提供了一个统一基准,给出了单模态和多模态方法的广泛基线结果,包括模态缺失和噪声条件下的分析。实验结果证明了多模态融合的好处,并揭示了现实世界噪声和低光条件下仍存在的挑战。通过发布InCarEmo,我们旨在为强大、可解释且以人为本的车内情感理解建立全面基础,促进更安全、更具同理心的人车交互。
英文摘要
Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction. However, existing public datasets for in-cabin affective computing are largely limited to visual modalities and rarely include conversational information, making it difficult to capture the linguistic and interactive cues underlying driver emotion. To address these gaps, we introduce InCarEmo, a multimodal dataset for in-cabin emotion recognition and driver state monitoring. InCarEmo integrates RGB and infrared video, in-cabin audio, and dialogue text collected from scripted in-cabin scenarios designed to simulate realistic driver behaviors, covering diverse lighting conditions and driving contexts. The dataset supports three primary tasks: 1) multimodal emotion recognition, 2) fatigue detection, and 3) distraction monitoring. In addition to the original Chinese data, we construct an auxiliary English benchmark to support preliminary cross-lingual evaluation. We provide a unified benchmark with extensive baseline results across unimodal and multimodal methods, including analyses under modality-missing and noise conditions. Experimental results demonstrate the benefits of multimodal fusion and reveal remaining challenges under real-world noise and low-light conditions. By releasing InCarEmo, we aim to establish a comprehensive foundation for robust, interpretable, and human-centric in-cabin affective understanding, promoting safer and more empathetic driver-vehicle interaction.