裁决式字幕生成:用于严格零样本图像字幕生成的多智能体对齐评分与共识蒸馏束仲裁
Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
浏览论文内容
中文总结 AI 辅助
该研究提出裁决式字幕生成多智能体框架,通过多阶段对齐评分与共识蒸馏重排序,在不微调IFCap的情况下,显著提升COCO等数据集的零样本图像字幕生成性能,且可跨数据集迁移。
中文摘要 AI 辅助
零样本图像字幕生成(ZIC)指在字幕生成器训练时不使用配对的图像-字幕监督,仅依赖纯文本语料库和冻结的预训练图像-文本评分器来描述图像。现有的检索增强方法仅在检索时对图像-文本对齐进行一次评分,之后仅根据语言模型的概率确定字幕生成器的自回归束,导致解码器无法获得进一步的视觉接地反馈。该领域进展已停滞,自2024年以来没有方法能超越严格设定下的最佳结果。我们提出裁决式字幕生成(Adjudicated Captioning),这是一种推理时的多智能体框架,在不改变IFCap字幕生成器的情况下,在多个检查点恢复接地反馈。首先,我们在输入端安装更强的冻结检索编码器;其次,在检索与解码之间插入冻结的交叉注意力验证器,将排名前9的检索结果重新排序为前5;最后,在输出束处附加一个学习型重排序器,其由多层感知器TriFuse与记忆注意力Transformer MemAttend配对组成,这是整个流程中仅有的学习型组件,二者通过三个冻结评分器的博达共识蒸馏进行自监督训练,不使用配对的图像-字幕标签和参考字幕。在归纳式标题协议下,使用在不相交的COCO Karpathy验证束上拟合的重排序器并冻结应用于测试,该框架在COCO Karpathy上达到CIDEr 117.6、SPICE 21.9,较IFCap的108.0和20.3提升了9.6个CIDEr点,比最强的合成图像增强方法NES的109.9高出7.7,且无需重新训练字幕生成器。无训练的固定融合基线达到CIDEr 115.8,因此9.6的CIDEr提升中有7.8来自非学习型架构干预,剩余1.8来自学习型重排序器。该方案无需重新训练字幕生成器即可在COCO外迁移:在Flickr30k Karpathy上CIDEr提升8.1,在NoCaps上整体提升5.7。
英文摘要
Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods score image-text alignment once, at retrieval, then commit the captioner's autoregressive beam under language-model probability alone, leaving the decoder without further visual grounding feedback. Progress has stalled, with no method improving on the strict-regime best since 2024. We propose Adjudicated Captioning, an inference-time multi-agent framework that restores grounding feedback at multiple checkpoints over an unchanged IFCap captioner. First, we install a stronger frozen Retrieval Encoder at the input. Second, between retrieval and decoding we insert a frozen Cross-Attention Verifier that re-ranks the top-9 retrievals to top-5. Third, at the output beam we attach a learned Reranker pairing TriFuse, a multilayer perceptron, with MemAttend, a memory-attended transformer, the pipeline's only learned components; both are trained self-supervised by Borda-consensus distillation across the three frozen scorers, using no paired image-caption labels and no reference captions. Under the inductive headline protocol, with rerankers fit on the disjoint COCO Karpathy validation beam and applied frozen to test, the framework reaches CIDEr 117.6 and SPICE 21.9 on COCO Karpathy, up from 108.0 and 20.3 for IFCap, a +9.6 CIDEr gain, and +7.7 above NES, the strongest synthetic-image-augmented method at 109.9, without retraining the captioner. A training-free fixed-fusion baseline reaches 115.8 CIDEr, so +7.8 of the +9.6 gain comes from the non-learned architectural intervention and the remaining +1.8 from the learned rerankers. The same recipe transfers off-COCO without captioner retraining: +8.1 CIDEr on Flickr30k Karpathy and +5.7 on NoCaps overall.
发表机构
- AI Platform OneNexus, OneMount(OneMount AI平台OneNexus)
- School of Electronic Engineering, Soongsil University(崇实大学电子工程学院)
- MoAdata(MoAdata公司)
- University of Economics Ho Chi Minh City (UEH)(胡志明市经济大学)
机构由 AI 辅助整理,请以论文原文为准。