发表机构
Central South University; Peking University(中南大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大模型幻觉问题,提出Savor训练框架,通过置信度校准和GRPO优化,在推理时基于不确定性重新审视视觉证据,有效降低幻觉并保持通用能力。
AI 中文摘要
多模态大语言模型(MLLMs)在视觉问答和图像描述方面取得了显著进展,但它们仍会生成关于图像中未实际存在的对象、属性或关系的流畅表述。许多补救方法要么在测试时修改解码过程(这会增加延迟),要么使用诸如DPO变体之类的偏好进行微调(这教会模型哪种答案更受偏好,但并未教会模型何时自身的答案不可靠)。我们认为,校准后的自我评估是缺失的关键信号。我们提出了Savor,一个训练框架,它(i)通过引入令牌和答案置信度来扩展输出模式,(ii)使用组相对策略优化(GRPO)目标来优化策略,该目标惩罚校准误差和不良的弃权(不执行)决策,以及(iii)在推理时利用学习到的置信度,仅在模型不确定时重新审视视觉证据。在POPE、HallusionBench、AMBER和MMHal-Bench上,使用两个近期骨干网络(InternVL3-8B和Qwen3-VL-8B)进行的实验表明,Savor在降低幻觉的同时,保持了在MME和MMBench上的通用能力,并且其期望校准误差低于DPO和解码基线。
英文摘要
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that are not grounded in the image. Many remedies either modify decoding at test time, which adds latency, or fine tune with preferences such as DPO variants, which teach which answer is preferred but not when the model's own answer is unreliable. We argue that calibrated self assessment is the missing signal. We introduce Savor, a training framework that (i) augments the output schema with token and answer confidence, (ii) optimises the policy with a Group Relative Policy Optimisation (GRPO) objective that penalises calibration error and poor abstention decisions, and (iii) uses the learned confidence at inference time to revisit visual evidence only when the model is uncertain. Experiments on POPE, HallusionBench, AMBER and MMHal-Bench across two recent backbones (InternVL3-8B and Qwen3-VL-8B) show that Savor reduces hallucination while preserving general capability on MME and MMBench, with lower Expected Calibration Error than DPO and decoding baselines.
Comments33rd International Conference on Neural Information Processing (ICONIP 2026)