基于骨骼的零样本时空动作定位:弱监督预训练方法
Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining
浏览论文内容
中文总结 AI 辅助
该研究提出基于骨骼的零样本时空动作定位方法,通过弱监督视觉-语言预训练与场景混合判别对比学习,解决标注限制问题,在四个公开数据集上验证了有效性。
中文摘要 AI 辅助
针对基于骨骼的零样本时空动作定位,我们提出一种新颖的预训练策略,以估计人物实例的未见动作,同时通过新目标动作克服训练的高标注成本,并利用大规模动作场景数据集进行预训练。具体而言,我们的方法称为骨骼-语言特征池切换(Skeleton-Language feature Pooling Switching),引入弱监督视觉-语言预训练机制,该机制将预训练阶段的池化内核(用于在视频级别聚合骨骼特征并与每个视频的已知动作文本嵌入对齐)过渡到推理阶段,无需通过目标动作训练即可计算实例级特征。此外,我们提出场景混合判别对比学习(Scene-Mixed Discriminative Contrastive Learning),通过多实例学习(MIL)框架区分组合场景中的实例级动作。在四个公开的时空动作定位与分类数据集上的实验表明,该方法有效解决了标注限制问题。
英文摘要
We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target actions and pretraining using large-scale action scenery datasets. Specifically, our approach, termed Skeleton-Language feature Pooling Switching, introduces a weakly-supervised vision-language pretraining mechanism. This mechanism transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without training via target actions. Furthermore, we propose Scene-Mixed Discriminative Contrastive Learning to distinguish actions at the instance level within the combined scene through the MIL framework. Our experiments on four public spatio-temporal action localization and classification datasets demonstrate that the proposed method effectively addresses annotation limitations.
发表机构
- Konica Minolta, Inc.(柯尼卡美能达公司)
- Keio University(庆应义塾大学)
- CyberAgent(CyberAgent公司)
- NVIDIA(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。