arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有限资源下自监督图像与视频预训练的对照研究

A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources

Brunó B. Englert, Gijs Dubbelman

arXiv 2608.13183首次发表:更新:

发表机构

Eindhoven University of Technology(埃因霍温理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对照研究有限资源下的图像与视频自监督预训练,发现DINOv2式预训练性能最优,结合其与VideoMAE可提升图像任务性能但降低部分视频任务性能,为资源有限场景的视觉模型训练提供参考。

AI 中文摘要

视觉基础模型是图像与视频理解的核心,但通常需要大量数据与计算资源。当前视觉基础模型预训练所需的规模可能难以持续或并无必要,若能以更少资源获得有效模型,将带来显著益处。为更好理解自监督学习(SSL)目标在资源约束下的表现,我们在匹配的数据、架构与计算预算下,对图像与视频SSL目标开展对照研究。我们对比对比式、重建式、特征预测式与扩散式目标,在多样的图像与视频理解任务中评估独立训练及联合训练的图像-视频SSL方案。结果显示,DINOv2式预训练在有限资源下始终提供最强整体性能;此外,将DINOv2与VideoMAE等视频SSL目标结合,可大幅提升图像分类与分割性能,但会降低视频跟踪与相机位姿估计性能,揭示语义与几何表示学习间的重要权衡。这些发现表明,在资源有限场景下结合图像与视频SSL目标是有益的,同时凸显了需改进方法以更好平衡语义、时间与几何监督。

英文摘要

Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining visual foundation models may be unsustainable or unnecessary, and significant benefits arise when effective models can be obtained with fewer resources. To better understand how self-supervised learning (SSL) objectives behave under resource constraints, we conduct a controlled study of image and video SSL objectives under matched data, architecture, and compute budgets. We compare contrastive, reconstruction, feature-prediction, and diffusion objectives and evaluate both standalone and jointly trained image-video SSL formulations across a diverse set of image and video understanding tasks. Our results show that DINOv2-style pretraining consistently provides the strongest overall performance under limited resources. Furthermore, combining DINOv2 with video SSL objectives such as VideoMAE substantially improves image classification and segmentation performance, but degrades video tracking and camera-pose estimation performance, revealing an important tradeoff between semantic and geometric representation learning. These findings suggest that combining image and video SSL objectives can be beneficial in resource-limited settings, while highlighting the need for improved methods that better balance semantic, temporal, and geometric supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑