发表机构
Korea University; KAIST(高丽大学; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对MLLMs的物体幻觉问题,提出C²-DPO方法,通过最大化上下文偏好增益降低幻觉率,使Qwen2-VL-Instruct-2B的幻觉率相对降低36%且不损害通用推理能力。
AI 中文摘要
多模态大语言模型(Multimodal large language models, MLLMs)已取得快速进展,但仍存在物体幻觉问题——生成看似合理却与视觉输入不符的错误描述。直接偏好优化(Direct Preference Optimization, DPO)通过训练模型偏好非幻觉响应而非幻觉响应来缓解该问题,近期研究进一步用相关上下文丰富偏好数据。然而,DPO是否真正利用了此类上下文仍不明确。为探究此问题,我们提出上下文偏好增益(Contextual Preference Gain, CPG)这一简单指标,用于衡量提供相关上下文时模型偏好的增强程度。研究发现,更高的CPG始终对应更低的幻觉率,但标准DPO及其变体仅表现出有限的CPG,表明它们未充分利用上下文信息,因此仍易出现幻觉。为解决该问题,我们提出上下文校准DPO(Context-Calibrated DPO, C²-DPO),在保留原始偏好排序的同时直接最大化CPG。在多个基准测试中,C²-DPO大幅降低了幻觉,且未损害通用推理能力:相对于Qwen2-VL-Instruct-2B的物体幻觉基准(Object HalBench),其幻觉率降低了36%。代码可在该https URL获取。
英文摘要
Multimodal large language models (MLLMs) have made rapid progress, yet they still exhibit object hallucination, generating plausible but incorrect descriptions that are inconsistent with the visual input. Direct Preference Optimization (DPO) mitigates this by training models to prefer non-hallucinated responses over hallucinated ones, and recent efforts further enrich the preference data with relevant context. However, it remains unclear whether DPO actually leverages such context. To investigate this, we propose Contextual Preference Gain (CPG), a simple metric that measures how much a model's preference strengthens when relevant context is provided. We find that higher CPG consistently corresponds to lower hallucination, yet standard DPO and its variants exhibit only limited CPG, indicating that they underutilize contextual information and thus remain prone to hallucination. To address this, we propose Context-Calibrated DPO (C$^2$-DPO), which directly maximizes CPG while preserving the original preference ordering. Across multiple benchmarks, C$^2$-DPO substantially reduces hallucination without compromising general reasoning, relatively reducing the Object HalBench hallucination rate of Qwen2-VL-Instruct-2B by 36%. Code is available at https://github.com/mlvlab/C2-DPO
CommentsAccepted at ECCV2026