arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FUSE:基于面部视频的帧统一压力估计

FUSE: Frame-Unified Stress Estimation from Facial Video

Stefanos Gkikas, Thomas Kassiotis, Yang Guo, Guangliang Li, Giorgos Giannakakis

arXiv 2608.10442首次发表:更新:

AI 中文总结

本研究提出FUSE框架,将面部视频完整帧融合为统一表示,采用非对称注意力架构处理,在58人数据集上测试,t=15时准确率达69.44%,证实无需时间窗口划分即可实现高效压力检测。

AI 中文摘要

从面部视频自动检测压力为非侵入式情感监测提供了可行路径,但现有基于视频的方法通常会将完整记录分解为短时间窗口后再进行分类。这种设计引入了关于窗口长度、重叠和聚合的额外选择,同时限制了对整个记录中时间信息的直接分析。本研究提出了FUSE(Frame-Unified Stress Estimation,帧统一压力估计),这是一种面部视频压力检测框架,可将完整记录作为单个输入进行处理,无需时间窗口划分或外部分段。该方法的核心操作是:不将记录划分为短片段,而是将所有帧融合为一个统一的二维表示,以此估计压力状态。这种统一通过将时间维度折叠到空间表示的通道维度来实现,生成的高维输入采用统一的非对称注意力架构处理。当时间步长t=1时,FUSE将完整的120秒记录作为单个输入,对应30fps下的3600帧。在包含58名受试者的压力数据集上,采用分层受试者级协议对7种时间步长配置(从全帧输入到稀疏子采样)进行评估。FUSE在t=15时达到最高测试准确率69.44%,而全帧配置仍具有竞争力,准确率为69.03%。在步长范围内,计算成本从12.48 GFLOPs到348.78 GFLOPs不等,显示出时间密度与效率之间的权衡。这些结果表明,在该场景下,有效的面部视频压力检测无需时间窗口划分,且可在单个统一架构中实现完整记录的推理。

英文摘要

Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approaches commonly decompose full recordings into short temporal windows before classification. This design introduces additional choices regarding window length, overlap, and aggregation, while limiting direct analysis of temporal information across the entire recording. In this study, we present FUSE (Frame-Unified Stress Estimation), a facial-video stress detection framework that processes complete recordings as a single input without temporal windowing or external segmentation. The name reflects the defining operation of the method: rather than dividing a recording into short clips, all frames are fused into one unified two-dimensional representation from which the stress state is estimated. This unification is realized by folding the temporal dimension into the channel dimension of the spatial representation, and the resulting high-dimensional input is processed using a unified asymmetric-attention architecture. At a temporal stride of t = 1, FUSE retains the full 120-second recording as one input, corresponding to 3,600 frames at 30 fps. Experiments on a 58-subject stress dataset using a stratified subject-level protocol evaluate seven temporal-stride configurations, ranging from full-frame input to sparse subsampling. FUSE achieves the highest test accuracy of 69.44% at t = 15, while the full-frame configuration remains competitive at 69.03%. Across the stride range, computational cost varies from 12.48 to 348.78 GFLOPs, showing the trade-off between temporal density and efficiency. These results demonstrate that temporal windowing is not required for effective facial-video stress detection in this setting, and that complete-recording inference can be achieved within a single unified architecture.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑