arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于微表情识别的基于动作单元引导的合成视频生成

AU-Guided Synthetic Video Generation for Micro-Expression Recognition

Pei-Sze Tan, Sailaja Rajanala, Yee-Fan Tan, Raphael C. -W. Phan, Huey-Fang Ong

arXiv 2607.10860首次发表:更新:

发表机构

CyPhi AI Lab, Monash University Malaysia(马来西亚莫纳什大学CyPhi人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有微表情识别数据集的局限,提出基于动作单元引导生成合成微表情数据集EquiME的方法,经实验其在跨数据集微表情识别任务中性能有竞争力,且在不同架构中变化低,为微表情识别研究提供了新资源。

AI 中文摘要

微表情识别受现有数据集规模小、人口统计覆盖范围窄和情感标签受限的限制。我们引入了EquiME,这是一个通过动作单元引导的图像到视频生成构建的合成微表情数据集。EquiME包含从15K源面部图像生成的75K个视频,涵盖五种目标情绪,以及自动推断的人口统计元数据和视频质量测量。我们使用帧对相似度、空间变化和无参考感知质量指标评估EquiME,并在SAMM和CASME II上进行跨数据集MER实验。在EquiME上训练的模型在SAMM和CASME II上实现了有竞争力的跨数据集性能,且在四种评估架构中变化相对较低。本文重点关注数据集设计、用于视频生成的结构化动作单元条件管道,以及将EquiME评估为合成MER资源所需的实证证据。

英文摘要

Micro-expression recognition is limited by the small scale, narrow demographic coverage, and restricted emotion labels of existing datasets. We introduce EquiME, a synthetic micro-expression dataset built from AU-guided image-to-video generation. EquiME contains 75K videos generated from 15K source face images across five target emotions, together with automatically inferred demographic metadata and video-quality measurements. We evaluate EquiME using frame-pair similarity, spatial variation, and no-reference perceptual-quality metrics, together with cross-dataset MER experiments on SAMM and CASME II. Models trained on EquiME achieve competitive cross-dataset performance on SAMM and CASME II and show comparatively low variation across the four evaluated architectures. This paper focuses on the dataset design, the structured AU-conditioning pipeline used for video generation, and the empirical evidence needed to assess EquiME as a synthetic MER resource. Project page: https://kirito-blade.github.io/me-vlm/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑