arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29553cs.CVcs.AI

UNWIND:无需时间窗口划分的任意长度面部视频压力检测

UNWIND: Any-Length Facial Video for Stress Detection without Temporal Windowing

Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez

首次发表
浏览论文内容

中文总结 AI 辅助

提出UNWIND框架,将完整面部视频作为单一输入,通过折叠时间维度至通道维度并采用不对称注意力架构,实现无需时间窗口划分的压力检测,在58名受试者数据集上达到70.02%的准确率。

中文摘要 AI 辅助

基于面部视频的自动压力识别为情感监测提供了一种非接触式方法。然而,现有的大多数基于视频的方法在分类前会将完整录像划分为较短的时间片段。这种分段处理需要额外决定片段时长、重叠率以及预测聚合方式,并可能限制模型利用分布在整个录像中的信息。我们提出了UNWIND,一种用于压力检测的面部视频框架,该框架将完整录像作为单一模型输入进行分析,无需时间窗口划分或外部分段。UNWIND通过将视频的时间维度折叠到二维空间表示的通道维度来重新组织视频,随后通过统一的不对称注意力架构进行处理。在时间步长$\ au=1$的情况下,该框架在单次输入中处理整个120秒序列,对应以30帧每秒采样的3,600帧。我们在包含58名受试者的压力数据集上评估了七种时间步长设置,采用分层受试者级别的协议,覆盖从密集帧保留到稀疏时间采样的配置。最高测试准确率为$70.02\%$,在$\ au=15$时获得,而在$\ au=1$时处理所有帧达到的准确率为$69.73\%$,两者相当。计算需求在评估的步长设置中从12.48到348.78 GFLOPs不等,展示了时间采样密度与计算效率之间的平衡。研究结果表明,无需将录像划分为时间窗口即可实现有效的面部视频压力识别,并且可以在单个统一模型内完成完整录像的推理。

英文摘要

Automatic stress recognition from facial video provides a non-contact approach for affective monitoring. However, most existing video-based methods divide complete recordings into shorter temporal segments before performing classification. Such segmentation requires additional decisions concerning segment duration, overlap, and prediction aggregation, and may restrict the model from exploiting information distributed across the entire recording. We introduce UNWIND, a facial-video framework for stress detection that analyzes a complete recording as a single model input, eliminating the need for temporal windowing or external segmentation. UNWIND reorganizes the video by folding its temporal dimension into the channel dimension of a two-dimensional spatial representation, which is subsequently processed through a unified asymmetric-attention architecture. With a temporal stride of $τ=1$, the framework processes the entire $120$-second sequence, corresponding to $3{,}600$ frames sampled at $30$~fps, in a single input. We evaluate seven temporal-stride settings on a stress dataset comprising $58$ subjects, using a stratified subject-level protocol that covers configurations from dense frame retention to sparse temporal sampling. The highest test accuracy, $70.02\%$, is obtained at $τ=15$, while processing all frames at $τ=1$ achieves a comparable accuracy of $69.73\%$. Computational requirements range from $12.48$ to $348.78$ GFLOPs across the evaluated stride settings, illustrating the balance between temporal sampling density and computational efficiency. The findings show that effective facial-video stress recognition can be achieved without dividing recordings into temporal windows and that complete-recording inference can be performed within a single unified model.

发表机构

  • Honda Research Institute Japan(本田日本研究院)
  • Hellenic Mediterranean University(希腊地中海大学)

机构由 AI 辅助整理,请以论文原文为准。

↑