发表机构
Rensselaer Polytechnic Institute(伦斯勒理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对自动驾驶片段分类需求,测试基于JEPA的无标签新颖性评分方法,发现其在跨数据集场景看似有效,实则是域偏移导致,单数据集基准下失效,瓶颈为自监督目标。
AI 中文摘要
现代自动驾驶车队记录的视频远超人工审核员可检查的数量,这催生了对自动片段分类机制的需求,以筛选出罕见且值得审核的片段,从而可对驾驶模型进行微调,使其更好地应对非理想场景。我们测试了一种无标签方法,该方法通过自监督联合嵌入预测架构(JEPA)的预测误差“新颖性”对片段进行评分:将冻结的V-JEPA视频编码器与轻量级预测头配对以重建被掩码的片段嵌入,嵌入难以预测的片段会被标记为有趣。在以一个数据集训练、对其他数据集的视频进行测试的现实协议下评估时,该方法看似非常有效。我们证明,这种表面上的成功实际上是域偏移的结果:在来自单个数据集的公平基准上,该机制的表现降至随机水平,与简单的无训练基线相当。对相同冻结嵌入进行轻度监督探测的平均精度几乎翻倍,表明瓶颈确实在于自监督目标,而非表示本身。我们将此作为一项关于评估自监督学习有效性的研究,其中跨数据集协议可能会暗中奖励域分离而非新颖性。
英文摘要
Modern autonomous-driving fleets record far more video than human reviewers can inspect. This motivates the need for an automatic clip triage mechanism, to surface rare and review-worthy clips, so that driving models can be fine-tuned to better handle unideal circumstances. We test a label-free approach that scores clips by the prediction-error "novelty" of a self-supervised joint-embedding predictive architecture (JEPA); a frozen V-JEPA video encoder is paired with a lightweight predictor head to reconstruct masked clip embeddings, and clips whose embeddings are hard to predict are flagged as interesting. Evaluated under a realistic protocol that trains on one dataset and tests against footage from others, this approach appears highly effective. We show that this apparent success is actually a domain-shift consequence: on a fair benchmark drawn from a single dataset, this mechanism collapses to chance and is on par with simple no-training baselines. A lightly supervised probe on the same frozen embeddings results in almost double the average precision, indicating that the bottleneck is indeed the self-supervised objective, rather than the representation. We present this as a study for evaluating the effectiveness of self-supervised learning, where cross-dataset protocols can silently reward domain separation over novelty.