arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

屏蔽关键内容:面向自动驾驶的显著性引导视频自监督学习

Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving

Christopher Lang, Alexander Braun, Abhinav Valada

arXiv 2608.17178首次发表:更新:

发表机构

Robert Bosch GmbH; University of Freiburg(罗伯特·博世有限公司; 弗赖堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视频自监督学习随机掩码削弱自动驾驶安全关键线索的问题,提出 V-JEPA4A 模型,采用显著性驱动掩码策略,在四个驾驶基准上取得优于 V-JEPA 的性能,仅增加少量预训练开销。

AI 中文摘要

通过掩码时空预测实现的视频自监督学习已成为从无标签数据中学习特征表示的有前景范式。然而,现有方法通常依赖随机掩码,其会不加区分地去除区域,不考虑区域的语义或时间相关性。在以自身为中心的驾驶视频中,这会削弱 pretext 信号,因为行人、车辆、车道边界和动态交互等安全关键线索通常仅占帧的一小部分,但对下游感知至关重要。我们引入 V-JEPA4A,这是 V-JEPA 针对自动驾驶的领域专用变体,其基于公开驾驶视频进行预训练,采用了新颖的显著性驱动掩码策略。该策略考虑语义和时间相关上下文,根据语义重要性和时间相关性保留并预测上下文,在保留掩码预测效率的同时,实现更具信息量的表示学习。我们在涵盖跟踪、语义分割和深度估计的四个驾驶基准上评估所得编码器。结果表明,V-JEPA4A 在 BDD100k MOT 上相比采用随机掩码的 V-JEPA 减少了 25% 的身份切换,在 Cityscapes 上达到 73.2 mIoU,在 KITTI-2015 深度估计上达到 3.75 RMSE,同时仅产生约 14% 的额外预训练迭代开销。

英文摘要

Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, existing methods typically rely on random masking, which indiscriminately removes regions irrespective of their semantic or temporal relevance. In ego-centric driving videos, this can weaken the pretext signal since safety-critical cues such as pedestrians, vehicles, lane boundaries, and dynamic interactions often occupy only a small portion of the frame, yet are central to downstream perception. We introduce V-JEPA4A, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy. It accounts for semantically and temporally relevant context. The proposed policy preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation learning while retaining the efficiency of masked prediction. We evaluate the resulting encoders on four driving benchmarks spanning tracking, semantic segmentation, and depth estimation. The results demonstrate that V-JEPA4A reduces identity switches on BDD100k MOT by 25% over V-JEPA with random masking, achieves 73.2 mIoU on Cityscapes, and 3.75 RMSE on KITTI-2015 depth, while incurring only ~14% additional pre-training iteration overhead.

CommentsAccepted at GCPR 2026. The final publication will be available through Springer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑