发表机构
Guangxi University; School of Computer, Electronics and Information(广西大学; 计算机与电子信息学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对零样本图像描述中合成监督的实体级错位问题,提出即插即用框架ReCap,通过实体级重对齐与自适应加权策略优化合成数据,在两类基准上实现最优性能。
AI 中文摘要
零样本图像描述旨在无需带注释的图像-文本对即可生成图像描述。近期方法利用文本到图像模型从纯文本语料库合成训练数据,但多数聚焦于提升整体数据质量。与之相反,我们发现合成图像-文本错位往往是结构化且细粒度的:配对可能在全局层面看似合理,却存在实体缺失或属性误对齐的问题,从而降低监督保真度。因此,基于全局相似度进行图像重匹配或再生的方法,虽可提升表面合理性,却无法系统性修复实体级错位。为解决该问题,我们提出ReCap,即一种即插即用框架,将合成数据优化从隐式全局匹配转向显式细粒度重对齐。具体而言,ReCap利用检测到的图像支持实体引导描述重写,以实现实体级对应,生成更忠实的合成监督。此外,我们引入自适应动态加权学习策略,在训练过程中降低不可靠合成配对的权重。作为通用框架,ReCap可集成到现有合成数据流程中。大量实验表明,ReCap在域内和跨域零样本图像描述基准上均持续提升图像-文本一致性,并达到了当前最优性能。
英文摘要
Zero-shot image captioning aims to generate image descriptions without annotated image-text pairs. Recent approaches exploit text-to-image models to synthesize training data from text-only corpora, but most focus on improving overall data quality. In contrast, we observe that synthetic image-text misalignment is often structured and fine-grained: pairs may remain globally plausible while containing missing entities or misgrounded attributes, thereby degrading supervision fidelity. As a result, methods based on global similarity for image rematching or regeneration may improve apparent plausibility, but cannot systematically repair entity-level misalignment. To address this issue, we propose ReCap, a plug-and-play framework that shifts synthetic data refinement from implicit global matching to explicit fine-grained realignment. Specifically, ReCap enforces entity-level correspondence by using detected image-supported entities to guide caption rewriting, yielding more faithful synthetic supervision. In addition, we introduce an adaptive dynamic weighted learning strategy to downweight unreliable synthetic pairs during training. As a general framework, ReCap can be integrated into existing synthetic-data pipelines. Extensive experiments show that ReCap consistently improves image-text consistency and achieves state-of-the-art performance on both in-domain and cross-domain zero-shot image captioning benchmarks.
CommentsAccepted to the 34th ACM International Conference on Multimedia (ACM MM 2026). 16 pages, 7 figures