arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于提示条件场景奖励的可控图像字幕生成

Controllable Image Captioning with Prompt-Conditioned Scene Rewards

Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim

arXiv 2609.00709首次发表:更新:

发表机构

POSTECH; Graduate School of Artificial Intelligence, POSTECH; Department of Computer Science and Engineering, POSTECH(浦项科技大学; 浦项科技大学人工智能研究生院; 浦项科技大学计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出FoCUS可控图像字幕方法,通过提示条件场景奖励实现细粒度语义控制,引入SCoPE基准评估,在不降低通用性能的情况下提升了字幕可控性与质量。

AI 中文摘要

大型视觉语言模型(VLM)能生成流畅的图像描述,但语义控制能力有限:用户无法可靠指定字幕应强调属性、关系还是特定图像区域。我们提出使用场景奖励的细粒度字幕控制方法(FoCUS),这是一种可控图像字幕生成方法,允许用户通过自然语言控制提示引导字幕朝向特定语义重点。核心思路是基于场景图对齐组件分数的提示条件控制目标,生成的字幕会被解析并对齐到场景图组件(如对象、属性和关系),这些组件会根据请求的重点被赋予不同权重,包括负权重。我们使用GRPO优化该目标,并通过更严格的对象有效性阈值以及基于推理的属性和关系评分验证进一步提升其可靠性。为评估可控性,我们引入语义控制与精度评估基准(SCoPE),该基准包含对比性的包含/避免约束,用于测量目标内容覆盖率和范围外内容的抑制效果。在两个VLM主干上的实验表明,FoCUS在不降低通用字幕性能的前提下,持续提升了可控性和细粒度字幕质量。

英文摘要

Large Vision-Language Models produce fluent image descriptions but offer limited semantic control: users cannot reliably specify whether captions should emphasize attributes, relations, or particular image regions. We present Fine-grained Captioning Control Using Scene Rewards (FoCUS), a controllable image captioning method that lets users steer captions toward specific semantic emphases through natural-language control prompts. The core idea is a prompt-conditioned control objective based on scene-graph-aligned component scores. Generated captions are parsed and aligned to scene-graph components such as objects, attributes, and relations. These components are differentially weighted, including negative weights, according to the requested emphasis. We optimize this objective with GRPO and further improve its reliability through a stricter object validity threshold and reasoning-based verification for attribute and relation scoring. To evaluate controllability, we introduce Semantic Control and Precision Evaluation (SCoPE), a benchmark with contrastive Include/Avoid constraints for measuring both target content coverage and out-of-scope suppression. Experiments on two VLM backbones show that FoCUS consistently improves controllability and fine-grained caption quality without degrading general caption performance.

CommentsEMNLP 2026 Main (26 pages); Project website: https://focus-emnlp2026.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑