发表机构
Beijing University of Posts and Telecommunications; The University of Hong Kong(北京邮电大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出受生物启发的框架,利用原始视频的运动边界生成伪实例掩码,监督单张图像编码器,在多视觉任务上实现优于基线的性能,为可扩展视觉预训练提供新方向。
AI 中文摘要
以物体为中心的视觉表征对物理世界感知至关重要,但现有视觉预训练方法常捕获语义类别,却未保留单个实例的身份与连贯性。我们提出一种受生物启发的框架,用于从原始视频中学习单张图像的以物体为中心的表征。该方法将运动边界作为物体级分组的来源:利用现成的光流和聚类生成伪实例掩码,通过像素级成对度量学习监督单张图像编码器。该框架既不需要人工标注,也不需要相机标定。我们首先从7163小时的驾驶和网络视频中获取1.95亿个伪标注帧,随后通过结合模型提案与运动证据的运动验证自监督训练,将监督范围扩展至4.21亿个帧。我们训练编码器至Swin-H,并将学习到的表征蒸馏到一系列Swin主干网络中。在单目深度估计、3D物体检测、3D占用预测以及端到端规划任务中,所得模型相较于监督和自监督预训练基线取得了具竞争力或更优的性能,在对几何和实例敏感的任务上迁移表现尤为突出。这些结果表明,源自运动的监督可教会静态图像编码器表征视觉实例,为可扩展的视觉预训练提供了互补方向。
英文摘要
Object-centric visual representations are important for physical-world perception, but existing visual pretraining methods often capture semantic categories without preserving the identity and coherence of individual instances. We present a biologically inspired framework that learns object-centric representations for single images from raw videos. Our approach uses motion boundaries as a source of object-level grouping: off-the-shelf optical flow and clustering produce pseudo-instance masks, which supervise a single-image encoder with pixel-level pairwise metric learning. The framework requires neither human annotations nor camera calibration. We first obtain 195 million pseudo-labeled frames from 7,163 hours of driving and web videos, then expand the supervision to 421 million frames with Motion-Verified Self-Training, which combines model proposals with motion evidence. We train encoders up to Swin-H and distill the learned representations into a family of Swin backbones. Across monocular depth estimation, 3D object detection, 3D occupancy prediction, and end-to-end planning, the resulting models achieve competitive or superior performance relative to supervised and self-supervised pretraining baselines, with particularly strong transfer on geometry- and instance-sensitive tasks. These results show that motion-derived supervision can teach static image encoders to represent visual instances, providing a complementary direction for scalable visual pretraining.