arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一段文本胜过千条标题:重新思考视觉-语言检索的文本监督

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

Mahyar Ghazanfari, Amin Tabrizian, Arsyi Aziz, Binshuai Wang, Peng Wei

arXiv 2608.05260首次发表:更新:

发表机构

The George Washington University(乔治华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-语言检索,通过系统对比单标题与段落文本监督,发现仅微调BLIP文本编码器的段落监督模型在DOCCI等任务上优于Long-CLIP,为文本粒度与检索性能的关系提供了新见解。

AI 中文摘要

CLIP和BLIP等对比视觉-语言模型通常基于短图像标题进行训练,这限制了它们从详细文本描述中检索图像的能力。尽管Long-CLIP等方法通过位置嵌入插值扩展了token限制,但我们提出一个更简单的问题:仅训练文本粒度是否决定长文本检索性能?我们针对对比图像-文本检索,对从单条标题到多句子段落的监督进行了系统研究。基于Qwen2-VL和Llama 3.2 Vision构建合成流水线,我们为50万张CC3M图像生成多样化标题、难负样本及经质量评分的段落。为隔离文本粒度的影响,我们仅微调BLIP文本编码器,同时固定视觉编码器,共设置10种训练配置。我们的段落监督模型在ShareGPT4V上可匹配Long-CLIP-L的性能,在DOCCI的图像到文本检索任务上超出其14个百分点以上,且无需架构改动。我们进一步表明,段落监督可实现长token序列的有效利用,而仅标题训练在token超过60时性能会下降;增加标题多样性可提升短标题检索效果但收益递减,而段落监督始终有益于长描述基准,且仅文本微调时难负样本会产生不利影响。在Flickr30k、COCO、ShareGPT4V和DOCCI上的评估,为文本粒度、检索方向与描述长度间的权衡提供了全面分析。

英文摘要

Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions. While methods such as Long-CLIP extend the token limit through positional embedding interpolation, we ask a simpler question: does training text granularity alone determine long-text retrieval performance? We present a systematic study of supervision ranging from single captions to multi-sentence paragraphs for contrastive image-text retrieval. Using a synthetic pipeline based on Qwen2-VL and Llama 3.2 Vision, we generate diverse captions, hard negatives, and quality-scored paragraphs for 500K CC3M images. To isolate the effect of text granularity, we fine-tune only the BLIP text encoder while keeping the vision encoder frozen across 10 training configurations. Our paragraph-supervised models match Long-CLIP-L on ShareGPT4V and outperform it by more than 14 points on DOCCI for image-to-text retrieval, without architectural changes. We further show that paragraph supervision enables effective use of long token sequences, whereas caption-only training degrades beyond 60 tokens. Increasing caption diversity improves short-caption retrieval with diminishing returns, while paragraph supervision consistently benefits long-description benchmarks and hard negatives prove detrimental in text-only fine-tuning. Evaluations on Flickr30k, COCO, ShareGPT4V, and DOCCI provide a comprehensive analysis of the trade-offs between text granularity, retrieval direction, and description length.

CommentsAccepted at the MUCG workshop, ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑