arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

什么使得合成困难负样本在视觉-语言预训练中有效?

What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining?

Nikos Giakoumoglou, Paschalis Giakoumoglou, Andreas Floros, Kleanthis Marios Papadopoulos, Tania Stathaki

arXiv 2610.09700首次发表:更新:

发表机构

Imperial College London; Centre for Research and Technology Hellas(伦敦帝国理工学院; 希腊研究与技术中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过几何分析揭示合成困难负样本在视觉-语言预训练中的两种失败模式,提出SNAP方法生成不涉及正样本的模态内困难负样本,在CLIP和FLIP上带来零样本检索、分类和线性探针的稳定提升。

AI 中文摘要

在表示空间中生成的合成困难负样本已被证明对单模态自监督学习有效,但将这一思想迁移到视觉-语言预训练并非易事。我们分析了六种表示空间合成策略,并识别出它们在迁移到视觉-语言预训练时的两种失败模式:跨模态构造产生过于简单的负样本或将它们拉向查询,以及模态内构造包含匹配的正样本。我们还观察到在使用合成困难负样本和可学习温度训练时出现logit尺度饱和现象,并发现固定温度能提升下游性能。基于这一几何分析,我们提出SNAP,它生成模态内困难负样本,且从不涉及任一模态的正样本,从而完全避免这两种失败模式。SNAP是模型无关的,不需要外部生成模型,且仅增加不到10%的训练时间开销。在CLIP和FLIP之上,跨多种架构和数据集进行评估,SNAP在零样本检索、零样本分类和线性探针评估上均带来一致改进。

英文摘要

Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to vision-language pretraining: cross-modal constructions that produce overly easy negatives or pull them toward the query, and intra-modal constructions that incorporate the matched positive. We also observe logit-scale saturation when training with synthetic hard negatives and a learnable temperature, and find that fixing the temperature improves downstream performance. Using this geometric analysis we propose SNAP, which generates intra-modal hard negatives that never involve the positive from either modality, avoiding both failure modes entirely. SNAP is model-agnostic, requires no external generative models, and adds less than 10% training time overhead. Evaluated on top of CLIP and FLIP across multiple architectures and datasets, SNAP delivers consistent improvements on zero-shot retrieval, zero-shot classification, and linear probe evaluation.

CommentsACCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑