使用双循环自动编码器进行时空面部动作单元检测以用于驾驶员监测
Spatiotemporal Facial Action Unit Detection using Twin Cycle Autoencoders for Driver Monitoring
浏览论文内容
中文总结 AI 辅助
研究针对驾驶场景中面部动作单元检测难题,提出双循环自动编码器TCA,由空间和时间循环自动编码器分支组成,通过特定损失耦合与融合,在多数据集上评估效果优于基线,能维持实时吞吐量,可用于生产级ADAS。
中文摘要 AI 辅助
驾驶员监测系统(DMS)越来越依赖面部线索实时推断困倦、分心和认知负荷。基于面部动作编码系统(FACS)的面部动作单元(AUs)能客观呈现这些状态,但在驾驶场景中自动检测因光照低且变化、部分遮挡、头部姿势变化以及相关AU激活的微妙性和短暂性而变得复杂。现有AU探测器大多分别处理空间外观和时间动态,限制了利用大量未标记驾驶视频的自监督信号的能力。我们提出双循环自动编码器(TCA),由两个耦合的循环一致自动编码器分支组成:空间循环自动编码器通过图像级循环一致性从身份中分离出与AU相关的外观,时间循环自动编码器在潜在AU轨迹上强制前后一致性以捕捉起始-顶点-偏移动态。两个分支通过跨分支潜在对齐损失耦合,并在多标签AU分类前通过注意力模块融合。我们在DISFA和BP4D基准以及车内自然驾驶数据集上评估TCA,观察到其相对于CNN-RNN、3D-CNN和基于图的AU基线有持续改进,尤其对于与疲劳(AU45、AU43)和打哈欠(AU26)相关的低强度和快速转变的AUs。我们还表明该模型在嵌入式Jetson Xavier NX平台上维持实时吞吐量,支持其在生产级高级驾驶员辅助系统(ADAS)中的使用。
英文摘要
Driver monitoring systems (DMS) increasingly rely on facial cues to infer drowsiness, distraction, and cognitive load in real time. Facial Action Units (AUs), grounded in the Facial Action Coding System (FACS), provide an objective and interpretable representation of such states, but their automatic detection in the driving context is complicated by low and variable illumination, partial occlusion, head-pose variation, and the subtlety and short duration of relevant AU activations. Existing AU detectors largely treat spatial appearance and temporal dynamics separately, limiting their ability to exploit self-supervisory signal from abundant unlabeled driving video. We propose the Twin Cycle Autoencoder (TCA), a spatiotemporal architecture composed of two coupled cycle-consistent autoencoder branches: a Spatial Cycle Autoencoder that disentangles AU-relevant appearance from identity through image-level cycle consistency, and a Temporal Cycle Autoencoder that enforces forward-backward consistency over latent AU trajectories to capture onset-apex-offset dynamics. The two branches are coupled through a cross-branch latent alignment loss and fused via an attention module before multi-label AU classification. We evaluate TCA on the DISFA and BP4D benchmarks and on an in-cabin naturalistic driving dataset, and observe consistent improvements over CNN-RNN, 3D-CNN, and graph-based AU baselines, particularly for low-intensity and rapidly transitioning AUs relevant to fatigue (AU45, AU43) and yawning (AU26). We further show the model sustains real-time throughput on an embedded Jetson Xavier NX platform, supporting its use in production-grade advanced driver assistance systems (ADAS).