基于对象(但类别无关)的视频域自适应
Object-based (yet Class-agnostic) Video Domain Adaptation
- UC Berkeley(加州大学伯克利分校)
- Tel Aviv University(特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出ODAPT框架,利用目标域中类别无关的对象稀疏标注帧实现视频动作识别模型的域自适应,在多个数据集上取得显著提升,并可与现有无监督方法结合进一步增强性能。
AI中文摘要:
现有的基于视频的动作识别系统通常需要密集标注,并且在与训练数据存在显著分布偏移的环境中表现不佳。当前的视频域自适应方法通常利用目标域数据子集上的全标注数据对模型进行微调,或者通过自举或对抗学习来对齐两个域的表示。受对象在近期监督式对象中心动作识别模型中的关键作用启发,我们提出了基于对象(但类别无关)的视频域自适应(ODAPT),这是一个简单而有效的框架,通过利用目标域中带有类别无关对象标注的稀疏帧集,将现有动作识别系统适配到新域。我们的模型在 Epic-Kitchens 中跨厨房适配时实现了 +6.5 的提升,在 Epic-Kitchens 与 EGTEA 数据集之间适配时实现了 +3.1 的提升。ODAPT 是一个通用框架,也可以与先前的无监督方法结合,当与自监督多模态方法 MMSADA 结合时在 Epic-Kitchens 上提供 +5.0 的提升,当添加到基于对抗的方法 TA$^3$N 上时提供 +1.7 的提升。
英文摘要:
Existing video-based action recognition systems typically require dense annotation and struggle in environments when there is significant distribution shift relative to the training data. Current methods for video domain adaptation typically fine-tune the model using fully annotated data on a subset of target domain data or align the representation of the two domains using bootstrapping or adversarial learning. Inspired by the pivotal role of objects in recent supervised object-centric action recognition models, we present Object-based (yet Class-agnostic) Video Domain Adaptation (ODAPT), a simple yet effective framework for adapting the existing action recognition systems to new domains by utilizing a sparse set of frames with class-agnostic object annotations in a target domain. Our model achieves a +6.5 increase when adapting across kitchens in Epic-Kitchens and a +3.1 increase adapting between Epic-Kitchens and the EGTEA dataset. ODAPT is a general framework that can also be combined with previous unsupervised methods, offering a +5.0 boost when combined with the self-supervised multi-modal method MMSADA and a +1.7 boost when added to the adversarial-based method TA$^3$N on Epic-Kitchens.