DEEPO:面向多模态大语言模型幻觉的双熵增强策略优化
DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs
- Zhejiang University(浙江大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态大语言模型强化学习中的幻觉问题,提出双熵增强策略优化(DEEPO),通过语义熵触发专家前缀和优势符号感知的Renyi预条件,恢复优势方差并校正自信错误,在VideoMMMU上显著提升性能并减少幻觉。
AI中文摘要:
强化学习(RL)被广泛用于提升多模态大语言模型(MLLMs)的推理能力,但其对幻觉的影响并不均匀。我们将此追溯到从奖励到参数更新的“校正链”中的两个薄弱环节。在采样层面,高难度查询(即具有高语义熵的查询)经常产生一致错误的样本组,使得组相对优势在幻觉风险最高的地方恰好坍缩为零。在优化层面,自信但错误的token对梯度不可见:分类策略的期望得分梯度范数随着其分布锐化而消失,因此最需要校正的预测获得最弱的更新。我们提出双熵增强策略优化(DEEPO),一种结合信号方差正则化与梯度预条件的双阶段增强方法:语义熵触发的专家前缀在高不确定性查询上注入有依据的延续,提供直接监督并恢复优势方差,而优势符号感知的Renyi预条件对抗logit级饱和,使校正能在操作置信度区间内到达自信的错误。两个分支单独均优于GRPO;它们的交互在VideoMMMU——我们评估套件中最复杂的长期任务(+4.0,95%置信区间[1.1, 6.9])——上具有统计显著性,并在其他任务上具有加性效果。DEEPO在保持准确性和训练稳定性的同时减少幻觉。
英文摘要:
Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At the rollout level, hard queries---those with high semantic entropy---frequently produce unanimously wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk is highest. At the optimization level, confident-but-wrong tokens are gradient-invisible: a categorical policy's expected score-gradient norm vanishes as its distribution sharpens, so the predictions that most need correction receive the weakest updates. We propose Dual-Entropy Enhanced Policy Optimization (DEEPO), a dual-stage enhancement combining signal variance regularization with gradient preconditioning: semantic-entropy-triggered expert prefixes inject grounded continuations on high-uncertainty queries, providing direct supervision and restoring advantage variance, while advantage-sign-aware Renyi preconditioning counteracts logit-level saturation so correction reaches confident errors in the operational confidence regime. Both branches improve over GRPO individually; their interaction is statistically significant on VideoMMMU---the most complex long-horizon task in our evaluation suite (+4.0$, 95\% CI [1.1, 6.9])---and additive elsewhere. DEEPO reduces hallucination while preserving accuracy and training stability.