arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HuC-VideoMAE:从合成数据中进行以人为中心的视频掩码自编码

HuC-VideoMAE: Human-Centric Video Masked Autoencoding from synthetic data

Ricardo Pizarro, Roberto Valle, José M. Buenaposada, Luis M. Bergasa, Luis Baumela

arXiv 2610.08433首次发表:更新:

发表机构

Universidad de Alcalá; Universidad Politécnica de Madrid; Universidad Rey Juan Carlos(阿尔卡拉大学; 马德里理工大学; 胡安卡洛斯国王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对合成数据上视频Transformer预训练性能不佳的问题,提出以人为中心的掩码方案,利用关键点和边界框,在NTU RGB+D和Toyota-Smarthome上显著优于标准VideoMAE,并缩小与Kinetics预训练差距的49%。

AI 中文摘要

现代动作识别模型依赖于在大量网络爬取视频(如Kinetics-700)上预训练的视频Transformer。然而,使用此类数据引发了伦理问题,因为通常未获得主体的同意。近期由运动捕捉数据生成的高质量合成视频数据集(如BEDLAM2.0)提供了一种有前景的伦理替代方案。在这项工作中,我们研究了在合成人体运动数据集上对视频Transformer进行自监督预训练。我们首先表明,直接应用标准的VideoMAE掩码策略会导致性能远低于在Kinetics上预训练。为解决这一局限,我们提出了一种以人为中心的掩码方案,该方案利用身体关键点和人物边界框区域。我们的方法鼓励模型在预训练期间专注于人体运动的结构和动态。在NTU RGB+D和Toyota-Smarthome上的实验表明,我们的方法显著优于在合成数据上的标准VideoMAE预训练,在NTU RGB+D跨视角-主体任务中,在不使用任何真实帧进行预训练的情况下,缩小了与Kinetics预训练差距的49%。为推广伦理动作识别模型的使用,我们将公开发布我们的预训练模型。

英文摘要

Modern action recognition models rely on video transformers pretrained on massive collections of web-crawled videos, such as Kinetics-700. However, the use of such data raises ethical concerns, as subjects' consent is typically not obtained. Recent high-quality synthetic video datasets generated from motion-capture data, such as BEDLAM2.0, offer a promising ethical alternative. In this work, we investigate self-supervised pretraining of video transformers on synthetic human-motion datasets. We first show that directly applying the standard VideoMAE masking strategy leads to substantially worse performance than pretraining on Kinetics. To address this limitation, we propose a human-centric masking scheme that leverages body keypoints and person bounding box regions. Our approach encourages the model to focus on the structure and dynamics of human motion during pretraining. Experiments on NTU RGB+D and Toyota-Smarthome demonstrate that our method significantly outperforms standard VideoMAE pretraining on synthetic data, closing 49% of the gap to Kinetics pretraining on NTU RGB+D cross-view-subject without using a single real frame during pretraining. To promote the use of ethical action recognition models, we will publicly release our pretrained models.

Journal refECCV 2026 Workshop on Privacy, Fairness, Accountability and Transparency in Computer Vision

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑