arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪些负样本重要?询问你的文本编码器:用于密集字幕检索的自适应相似度间隔

Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoying Tang

arXiv 2608.18521首次发表:更新:

发表机构

Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(深圳未来智能网络研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对密集字幕检索中InfoNCE目标函数过早饱和的问题,提出HN-CLIP方法,通过文本编码器的文本-文本几何构建自适应相似度间隔,在四个基准上提升R@1,且训练速度更快,仅用20%数据即可达最强基线。

AI 中文摘要

密集字幕检索近期通过在对比微调中引入分割、边缘图、大语言模型(LLM)过滤的字幕及跨模态模块得到改进,但这些方法大多沿用相同的InfoNCE目标函数,在强预训练初始化下其优化会过早饱和:在密集字幕任务中,该损失在首个epoch内80%的批次中降至10⁻³以下,且在47%的测量中其梯度在fp32下精确为零。我们发现该行为与密集字幕基准中大量近重复字幕密切相关,在多数易区分的负样本已被分离后,仍有少量高度相似的负样本未被解决。作为解决方案,我们提出HN-CLIP,利用文本编码器自身的文本-文本几何结构为每个负样本构建自适应相似度间隔。具体而言,将一个分离的字幕相似度矩阵添加到负对数几率中,无需挖掘、合成或重采样负样本,即可为更相似的字幕分配更大间隔。所得目标函数在训练期间仅需一个字幕相似度矩阵和掩码对数几率加法,无辅助数据、额外参数、离线预处理或推理时开销。在四个密集字幕检索基准上的大量实验表明,HN-CLIP比最强竞争对手的R@1提升了+2.4至+4.3,训练速度比GOAL快2.4倍,比StructXLIP快5.4倍。此外,所提目标函数在域内基准上改进了所有六个测试的微调框架,且仅用20%的训练数据即可达到最强的全数据基线。

英文摘要

Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, these methods largely inherit the same InfoNCE objective, whose optimization can prematurely saturate under a strong pre-trained initialization: on dense captions, the loss falls below 10^-3 on 80% of batches within the first epoch, while its gradient becomes numerically zero in 47% of measurements. We find that this behavior is closely related to the large number of near-duplicate captions in dense-caption benchmarks, where a few highly similar negatives remain unresolved after the easy majority has already been separated. As a remedy, we introduce HN-CLIP, which uses the text encoder's own text-text geometry to construct per-negative adaptive similarity margins. Specifically, a detached caption-similarity matrix is added to the negative logits, assigning larger margins to more similar captions without mining, synthesizing, or resampling negatives. The resulting objective requires only one caption-similarity matrix and a masked logit addition during training, with no auxiliary data, additional parameters, offline preprocessing, or inference-time overhead. Extensive experiments on four dense-caption retrieval benchmarks show that HN-CLIP improves over the strongest competitors by +2.5 to +4.0 R@1 while training 2.4x faster than GOAL and 5.4x faster than StructXLIP. Moreover, the proposed objective improves all six tested fine-tuning frameworks on the in-domain benchmarks and reaches the strongest full-data baseline with only 20% of the training data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑