arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13712cs.CVcs.AIcs.CLcs.MM

Groc-PO:用于真实多模态大语言模型的基于上下文的偏好优化

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs

Zhixiao Zheng, Zheren Fu, Zhiyuan Yao, Chunxiao Liu, Dongming Zhang, Zhendong Mao

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多模态大语言模型的不真实问题,提出基于上下文的偏好优化框架Groc-PO,构建相关数据集,通过多阶段偏好样本捕捉基础上下文,加强上下文相关推理,减轻跨阶段错误传播,提升模型性能。

中文摘要 AI 辅助

尽管多模态大语言模型取得了快速进展,但仍存在诸如视觉幻觉、内容编造和推理不准确等不真实问题,削弱了其可靠性和实用性。基于人类偏好的对齐方法如直接偏好优化(DPO)被广泛采用,但多模态推理错误常跨阶段传播,最终答案错误常源于早期基础阶段的错误,而标准DPO通常在最终答案层面进行偏好优化。为解决此问题,我们提出了用于多模态大语言模型的基于上下文的偏好优化框架Groc-PO。我们还构建了基于上下文的偏好数据集(GCPD),围绕对象基础、上下文基础和基于基础的推理三个阶段组织多阶段偏好样本,以捕捉基础上下文的形成、整合和利用。通过在多个基础阶段引入更明确的偏好监督,Groc-PO加强了上下文相关推理并减轻了跨阶段错误传播。大量实验表明,与标准DPO和其他强大基线相比,Groc-PO在减轻幻觉、忠实推理和整体可靠性方面取得了更好的性能,支持了更明确的基础监督对可信多模态推理的价值。

英文摘要

Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual hallucinations, content fabrication, and unfaithful reasoning, which substantially undermine their faithfulness and practical utility. Alignment methods based on human preference, such as Direct Preference Optimization (DPO), have been widely adopted to address these issues. However, multimodal reasoning errors often propagate across stages, and final-answer errors can often be traced to mistakes in early grounding stages, yet standard DPO typically applies preference optimization at the final-answer level. This credit-assignment challenge means that supervision for early grounding stages is indirect rather than stage-specific, making it difficult to suppress error propagation arising from grounding drift and context inconsistency. To address this, we propose Grounded Context Preference Optimization (Groc-PO), a grounded preference optimization framework for MLLMs. We further construct the Grounded Context Preference Dataset (GCPD), organizing multi-stage preference samples around three stages of Object Grounding, Contextual Grounding, and Grounded Reasoning, to capture the formation, integration, and utilization of grounded context. By introducing more explicit preference supervision over multiple grounded stages, Groc-PO strengthens context-dependent reasoning and mitigates cross-stage error propagation. Extensive experiments show that, compared with standard DPO and other strong baselines, Groc-PO achieves improved performance in hallucination mitigation, faithful reasoning, and overall reliability, supporting the value of more explicit grounded supervision for trustworthy multimodal reasoning.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Xiaomi Corporation(小米公司)
  • State Key Laboratory of Communication Content Cognition, People’s Daily Online(人民日报社传播内容认知国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑