arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成不确定性作为语义相似性学习的自监督信号

Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning

Enrico Pallotta, Sina Raoufi, Lars Doorenbos, Gianni Franchi, Juergen Gall

arXiv 2609.35341首次发表:更新:

发表机构

University of Bonn; Lamarr Institute for ML & AI; ENSTA Paris, Institut Polytechnique de Paris(波恩大学; 拉马尔机器学习与人工智能研究所; 巴黎高等先进技术学院,巴黎综合理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出利用文本到视频扩散模型的生成不确定性作为自监督信号,通过掩码学习稳定语义特征,无需人工标注即可提升视频语义相似性,实验验证其优于预训练特征。

AI 中文摘要

评估视频之间的语义相似性是计算机视觉中的一个基本挑战,对于从分布外(OOD)检测到视频检索等任务至关重要。然而,由于视频复杂的时空特性,定义和标注视频相似性众所周知地困难且昂贵。在本文中,我们提出了一种新颖的自监督方法,利用文本到视频(T2V)扩散模型中的生成不确定性,在无需人工标注的情况下学习语义相似性。我们的方法基于以下观察:T2V模型对于熟悉的概念会产生一致的输出,但在提示专业概念时表现出高方差和不确定性。我们利用这一行为来识别现有预训练表示(如VideoMAE和V-JEPA)中的稳定语义特征。具体而言,我们使用纯生成的数据学习这些嵌入上的掩码,鼓励模型保留在一般概念生成中保持一致的特性,同时丢弃与生成噪声或不确定性相关的特性。在三个关键任务上的实验结果表明,我们学习到的特征子空间始终优于原始预训练特征和基线特征选择方法。

英文摘要

Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this paper, we propose a novel self-supervised approach that leverages generative uncertainty from text-to-video (T2V) diffusion models to learn semantic similarity without human annotations. Our method is based on the observation that T2V models produce consistent outputs for familiar concepts but exhibit high variance and uncertainty when prompted with specialized concepts. We utilize this behavior to identify stable semantic features within existing pretrained representations, such as VideoMAE and V-JEPA. Specifically, we learn a mask over these embeddings using purely generated data, encouraging the model to retain features that remain consistent across generations of general concepts while discarding those associated with generative noise or uncertainty. Experimental results across three key tasks demonstrate that our learned feature subspaces consistently outperform original pretrained features and baseline feature selection methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑