发表机构
Adobe Research(奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对现有图像字幕模型遗漏视觉细节、多阶段系统延迟高的问题,提出SimLoss,实现单遍细粒度图像字幕生成,相关变体在精度、召回率等指标上表现优异且推理速度快。
AI 中文摘要
一幅图像价值千言,但大多数字幕生成模型仅用寥寥数语描述它。现代视觉-语言模型能生成流畅的高层级字幕,却常遗漏让图像具备视觉特异性的属性、计数、纹理、材质及空间关系。近期的多阶段系统通过生成、分解、验证与重写恢复部分此类细节,但代价是推理延迟显著升高。我们提出SimLoss,一种用于单遍细粒度图像字幕生成的无参考嵌入空间目标函数。SimLoss训练视觉-语言模型将其投影的隐态表示与冻结的图像嵌入通过InfoNCE对比损失对齐,在解码任何文本前提供密集视觉监督信号,且无需人工编写的细粒度字幕或多阶段流水线生成的伪字幕。我们将其实例化为SimLoss FFT(通过局部可用的嵌入模型反向传播)和SimLoss GRPO(将该模型视为黑盒奖励)。与单遍、多阶段验证、奖励优化及感知基准相比,全可微微调变体SimLoss FFT取得最高精度,同时几乎达到多阶段方法的F1分数,且保持单遍推理,运行速度约为多阶段流水线的20倍;基于奖励的变体SimLoss GRPO获得最强召回率。这些结果共同表明,嵌入空间监督可在单遍字幕生成器的延迟下恢复多阶段验证的质量。
英文摘要
An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level captions, yet routinely miss the attributes, counts, textures, materials, and spatial relations that make an image visually specific. Recent multi-stage systems recover some of these details through generation, decomposition, verification, and rewriting, but they do so at the expense of substantially higher inference latency. We propose SimLoss, a reference-free embedding-space objective for single-pass fine-grained image captioning. SimLoss trains a vision-language model to align its projected hidden-state representation with a frozen image embedding through an InfoNCE contrastive loss, supplying a dense visual supervision signal before any text is decoded, and requiring neither human-written fine-grained captions nor pseudo-captions from a multi-stage pipeline. We instantiate it as SimLoss FFT, which backpropagates through a locally available embedding model, and SimLoss GRPO, which treats that model as a black-box reward. Compared with single-pass, multi-stage verification, reward-optimized, and perception-aware baselines, the fully differentiable fine-tuning variant, SimLoss FFT, achieves the highest precision while nearly matching the F1 score of the multi-stage method, all while retaining single-pass inference and running roughly 20 times faster than the multi-stage pipeline. The reward-based variant SimLoss GRPO attains the strongest recall. Together, these results show that embedding-space supervision can recover the quality of multi-stage verification at the latency of a single-pass captioner.