arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.20994cs.CVcs.AIcs.LG

VESSA:基于视频的以对象为中心的自监督适应方法用于视觉基础模型

VESSA: Video-based objEct-centric Self-Supervised Adaptation for Visual Foundation Models

  • Departamento de Ciência da Computação, Universidade Federal de Minas Gerais (UFMG)(巴西联邦大学矿务学院计算机科学系(UFMG))
  • Recod.ai, Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(Recod.ai,计算机学院,坎皮纳斯州立大学(UNICAMP))
  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

Jesimon Barreto, Carlos Caetano, André Araujo, William Robson Schwartz

更新

AI总结:

VESSA提出一种基于视频的以对象为中心的自监督适应方法,利用短时多视角视频和自蒸馏,无需标注即可适应新领域,在分类任务上优于基础模型和先前方法。

AI中文摘要:

基础模型通过大规模预训练和监督微调,在多种任务上展现了强大的性能,从而推动了计算机视觉的发展。然而,在存在分布偏移和标签稀缺的领域,它们可能表现不佳,此时监督微调可能不可行。虽然持续的自监督学习用于模型适应在生成式语言模型中很常见,但这一策略已被证明对以视觉为中心的编码器模型效果不佳。为了应对这一挑战,我们提出了一种新的视觉基础模型自监督微调公式,该公式使模型无需标注即可适应新领域,仅利用短时多视角以对象为中心的视频。我们的方法称为VESSA:基于视频的以对象为中心的自监督适应方法,用于视觉基础模型。VESSA的训练技术基于自蒸馏范式,其中关键是要仔细调整预测头并部署参数高效的适应技术——否则,模型可能迅速遗忘其预训练知识并达到退化状态。VESSA显著受益于从以对象为中心的视频中不同帧获取的多视角对象观测,无需标注即可有效学习对多种采集条件的鲁棒性。通过在2个数据集上使用3个视觉基础模型进行的全面实验,VESSA在下游分类任务中相比基础模型和先前的适应方法展示了一致的改进。代码公开于https://github.com/jesimonbarreto/VESSA。

英文摘要:

Foundation models have advanced computer vision by enabling strong performance across diverse tasks through large-scale pretraining and supervised fine-tuning. However, they may underperform in domains with distribution shifts and scarce labels, where supervised fine-tuning may be infeasible. While continued self-supervised learning for model adaptation is common for generative language models, this strategy has not proven effective for vision-centric encoder models. To address this challenge, we introduce a novel formulation of self-supervised fine-tuning for vision foundation models, where the model is adapted to a new domain without requiring annotations, leveraging only short multi-view object-centric videos. Our method is referred to as VESSA: Video-based objEct-centric Self-Supervised Adaptation for visual foundation models. VESSA's training technique is based on a self-distillation paradigm, where it is critical to carefully tune prediction heads and deploy parameter-efficient adaptation techniques - otherwise, the model may quickly forget its pretrained knowledge and reach a degraded state. VESSA benefits significantly from multi-view object observations sourced from different frames in an object-centric video, efficiently learning robustness to varied capture conditions, without the need of annotations. Through comprehensive experiments with 3 vision foundation models on 2 datasets, VESSA demonstrates consistent improvements in downstream classification tasks, compared to the base models and previous adaptation methods. Code is publicly available at https://github.com/jesimonbarreto/VESSA.

补充信息

↑