发表机构
The George Washington University(乔治华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视觉-语言检索,通过系统对比单标题与段落文本监督,发现仅微调BLIP文本编码器的段落监督模型在DOCCI等任务上优于Long-CLIP,为文本粒度与检索性能的关系提供了新见解。
AI 中文摘要
CLIP和BLIP等对比视觉-语言模型通常基于短图像标题进行训练,这限制了它们从详细文本描述中检索图像的能力。尽管Long-CLIP等方法通过位置嵌入插值扩展了token限制,但我们提出一个更简单的问题:仅训练文本粒度是否决定长文本检索性能?我们针对对比图像-文本检索,对从单条标题到多句子段落的监督进行了系统研究。基于Qwen2-VL和Llama 3.2 Vision构建合成流水线,我们为50万张CC3M图像生成多样化标题、难负样本及经质量评分的段落。为隔离文本粒度的影响,我们仅微调BLIP文本编码器,同时固定视觉编码器,共设置10种训练配置。我们的段落监督模型在ShareGPT4V上可匹配Long-CLIP-L的性能,在DOCCI的图像到文本检索任务上超出其14个百分点以上,且无需架构改动。我们进一步表明,段落监督可实现长token序列的有效利用,而仅标题训练在token超过60时性能会下降;增加标题多样性可提升短标题检索效果但收益递减,而段落监督始终有益于长描述基准,且仅文本微调时难负样本会产生不利影响。在Flickr30k、COCO、ShareGPT4V和DOCCI上的评估,为文本粒度、检索方向与描述长度间的权衡提供了全面分析。
英文摘要
Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions. While methods such as Long-CLIP extend the token limit through positional embedding interpolation, we ask a simpler question: does training text granularity alone determine long-text retrieval performance? We present a systematic study of supervision ranging from single captions to multi-sentence paragraphs for contrastive image-text retrieval. Using a synthetic pipeline based on Qwen2-VL and Llama 3.2 Vision, we generate diverse captions, hard negatives, and quality-scored paragraphs for 500K CC3M images. To isolate the effect of text granularity, we fine-tune only the BLIP text encoder while keeping the vision encoder frozen across 10 training configurations. Our paragraph-supervised models match Long-CLIP-L on ShareGPT4V and outperform it by more than 14 points on DOCCI for image-to-text retrieval, without architectural changes. We further show that paragraph supervision enables effective use of long token sequences, whereas caption-only training degrades beyond 60 tokens. Increasing caption diversity improves short-caption retrieval with diminishing returns, while paragraph supervision consistently benefits long-description benchmarks and hard negatives prove detrimental in text-only fine-tuning. Evaluations on Flickr30k, COCO, ShareGPT4V, and DOCCI provide a comprehensive analysis of the trade-offs between text granularity, retrieval direction, and description length.
CommentsAccepted at the MUCG workshop, ECCV 2026