arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越少说:面向信息丰富且忠实的视觉-语言模型的细粒度对齐

Beyond Saying Less: Fine-Grained Alignment for Informative and Faithful Vision-Language Models

Xingming Long, Jie Zhang, Yuecong Min, Shiguang Shan, Xilin Chen

arXiv 2609.35294首次发表:更新:

发表机构

State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Zhongguancun Academy(中国科学院计算技术研究所; 中国科学院大学; 中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉-语言模型的对象幻觉问题,提出细粒度对齐框架,结合密集奖励与子句级信用分配,减少少说捷径,提升生成忠实性和判别性能。

AI 中文摘要

对象幻觉仍然是大型视觉-语言模型面临的主要挑战。虽然离策略偏好优化被证明是一种有效的解决方案,但在线策略强化学习提供了一个更有前景的方向,因为它直接针对模型当前的失败模式。然而,我们发现,如果没有细粒度的奖励制定和分配,在线策略优化往往会陷入一个简单的捷径:通过少说——即减少有效声明——来减少幻觉。为了全面解决这一问题,我们提出了一个细粒度对齐框架,该框架在数据层面将密集奖励信号与算法层面的精确信用分配相结合。具体而言,我们首先构建了密集对象存在与缺失(DOPA)数据集,以解决阻碍有效对象声明被验证和奖励的稀疏标注问题。DOPA在扩展词汇表上详尽地标注了每个概念的确定性存在与缺失,显著增加了在线策略 rollout 期间可靠奖励信号的密度。其次,我们提出了用于在线策略优化的子句级信用分配(SCAPO),以防止响应级共享优势允许局部幻觉损害同一响应内的所有其他有效输出。通过根据对象声明独立地为每个子句分配信用,SCAPO可以精确地强化忠实生成并惩罚幻觉。此外,我们利用由此产生的忠实图像描述作为辅助上下文,将生成收益迁移到判别任务。实验表明,我们的方法在生成任务中产生高度信息丰富且忠实的描述,同时在判别评估中带来明显的性能提升。

英文摘要

Object hallucination remains a major challenge for large vision-language models. While off-policy preference optimization proves to be an effective solution, on-policy reinforcement learning provides a more promising direction as it directly targets a model's current failure modes. However, we find that without fine-grained reward formulation and allocation, on-policy optimization often falls into an easy shortcut: reducing hallucinations merely by saying less---making fewer valid claims. To comprehensively resolve this, we propose a fine-grained alignment framework that couples dense reward signals at the data level with precise credit assignment at the algorithmic level. Specifically, we first construct the Dense Object Presence and Absence (DOPA) dataset to address sparse annotations that prevent valid object claims from being verified and rewarded. DOPA exhaustively annotates the deterministic presence and absence of every concept across an expanded vocabulary, significantly increasing the density of reliable reward signals during on-policy rollouts. Second, we propose Subsentence-level Credit Assignment for on-Policy Optimization (SCAPO) to prevent response-level shared advantages from allowing local hallucinations to compromise all other valid outputs within the same response. By assigning credit to each subsentence independently based on its object claims, SCAPO can precisely reinforce faithful generations and penalize hallucinations. Furthermore, we leverage the resulting faithful image descriptions as auxiliary context to transfer generative gains to discriminative tasks. Experiments demonstrate that our method produces highly informative, faithful descriptions in generative tasks while yielding clear performance gains on discriminative evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑