arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PatchGate:在冻结的视觉-语言模型中利用固有物体清单缩小语言表述差距

PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models

Jihyung Ko, Eunji Jung, Hyeongsub Kim, Ziseok Lee, Jae Won Cho, Sanghyun Jo, Kyungsu Kim

arXiv 2608.21819首次发表:更新:

发表机构

Seoul National University; LG CNS; Konkuk University; OGQ(首尔大学; LG CNS; 建国大学; OGQ)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对冻结VLMs图像字幕的物体表述问题,提出无训练框架PatchGate,通过VEX和VIED两个阶段提升可见物体覆盖率并减少物体幻觉,在AMBER数据集上取得了显著效果。

AI 中文摘要

视觉-语言模型(VLMs)中可靠的图像字幕生成要求字幕既精确又完整,既要避免提及无依据的物体,又要覆盖可见物体。现有的无训练方法主要解决前一项要求,通过在生成过程中干预模型预测的提及来抑制无依据的物体词,但由于这些方法仅作用于模型已可能提及的物体,输出中遗漏的可见物体仍难以恢复。本文提出PatchGate,一种无训练框架,在生成前从冻结的VLM中提取无提示的固有物体证据,用于缩小固有物体集与最终物体提及之间的差距。第一阶段为视觉证据提取(VEX),从语言模型解码器层的后半部分读取补丁级词汇证据,构建无任何任务提示的图像条件物体集;第二阶段为视觉证据包含-排除解码(VIED),利用该物体证据校准解码逻辑,促进有证据支持但表述不足的物体,抑制支持度弱但表述过度的物体。在AMBER数据集上,PatchGate无需外部检测器或微调,仅需一次额外的前向传播,即可提升物体级可靠性的两个方面:可见物体覆盖率从49.4提升至56.0(+13.4%),通过降低CHAIR指标将物体幻觉从7.5降至6.6(-12.0%)。

英文摘要

Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. Existing training-free methods primarily address the former requirement, suppressing unsupported object words by intervening on model-predicted mentions during generation. Because they operate only on objects the model is already likely to mention, visible objects omitted from the output remain difficult to recover. We propose PatchGate, a training-free framework that extracts prompt-free object evidence intrinsic to a frozen VLM before generation and uses it to narrow the gap between an intrinsic object set and final object mentions. In the first stage, Visual Evidence eXtraction (VEX) reads patch-level lexical evidence from the latter half of LM decoder layers and constructs an image-conditioned object set without any task prompt. In the second stage, Visual-Evidence Inclusion-Exclusion Decoding (VIED) uses this object evidence to calibrate decoding logits, promoting evidence-supported but under-verbalized objects and suppressing weakly supported but over-verbalized objects. On AMBER, PatchGate improves both sides of object-level reliability, increasing visible-object coverage from 49.4 to 56.0 (+13.4%) and reducing object hallucination by lowering CHAIR from 7.5 to 6.6 (-12.0%), without external detectors or fine-tuning and with one extra forward pass.

Comments31 pages, 9 figures. Code will be available

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑