arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04247eess.AScs.SD

CAD:用于减轻全模态大语言模型跨模态幻觉的冲突感知解码

CAD: Conflict-Aware Decoding to Mitigate Cross-Modal Hallucinations in Omnimodal Large Language Models

Yuchen Deng, Chang Sun, Hai-Tao Zheng, Feidiao Yang, Yuxing Han

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对全模态大语言模型的跨模态幻觉问题,提出无训练框架CAD,通过PCME和CAA评估冲突并调整解码权重,在多个基准数据集上较基线方法实现显著性能提升。

中文摘要 AI 辅助

全模态大语言模型(Omni-LLMs)集成了音频、视频和文本模态,但仍易受跨模态幻觉影响,即某一模态会不当影响对另一模态的预测。现有无训练解码器通过扰动或相关性加权来调节模态影响,但未评估联合视听分支内的预测兼容性。由于联合分支的差异可能表明有害干扰或有用互补性,可靠干预需同时评估差异幅度和可操作性。为此,我们提出无训练框架冲突感知解码(Conflict-Aware Decoding, CAD),包含潜在冲突幅度估计(Potential Conflict Magnitude Estimation, PCME)和冲突可操作性评估(Conflict Actionability Assessment, CAA)。PCME利用视听分歧以及联合预测与相关性加权单模态参考的偏差来量化潜在冲突;CAA则使用Dempster-Shafer可靠性折扣处理任务空间答案关系,通过查询相关性和答案决定性确定是否需要干预。当识别出可操作冲突时,CAD会选择性地将解码权重从联合分支重新分配至单模态分支。在CMM、AVHBench、WorldSense和VideoMME上的实验表明,CAD在多个视听主干网络上始终优于基础解码器和有竞争力的无训练方法;在Qwen2.5-Omni-7B上,CAD无需模型重新训练,分别在CMM和AVHBench上将整体准确率提升14.1和8.0个百分点。

英文摘要

Omnimodal large language models (Omni-LLMs) integrate audio, video, and text, yet remain vulnerable to cross-modal hallucinations, where one modality improperly influences predictions about another. Existing training-free decoders modulate modality influence through perturbation or relevance weighting, but do not assess predictive compatibility within the joint audio-visual branch. Because joint-branch discrepancies may indicate either harmful interference or useful complementarity, reliable intervention requires assessing both discrepancy magnitude and actionability. To this end, we propose Conflict-Aware Decoding (CAD), a training-free framework comprising Potential Conflict Magnitude Estimation (PCME) and Conflict Actionability Assessment (CAA). PCME quantifies potential conflict using audio-video disagreement and the deviation of the joint prediction from a relevance-weighted unimodal reference. CAA then applies Dempster-Shafer reliability discounting to task-space answer relations, using query relevance and answer decisiveness to determine whether intervention is warranted. When an actionable conflict is identified, CAD selectively reallocates decoding weight from the joint branch to the unimodal branches. Experiments on CMM, AVHBench, WorldSense, and VideoMME show that CAD consistently outperforms the base decoder and competitive training-free methods across multiple audio-visual backbones. On Qwen2.5-Omni-7B, CAD improves overall accuracy by 14.1 and 8.0 percentage points on CMM and AVHBench, respectively, without model retraining.

发表机构

  • Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
  • Pengcheng Laboratory(鹏城实验室)
  • School of Cyber Science and Engineering, Zhengzhou University(郑州大学网络空间安全学院)

机构由 AI 辅助整理,请以论文原文为准。

↑