arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DICA:多模态大语言模型中的双指标引导对比对齐

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models

Hao Yang, Jin Wang, Xuejie Zhang

arXiv 2607.23944首次发表:更新:

AI 中文总结

研究针对多模态大语言模型的问题,提出双指标引导对比对齐方法(DICA),通过跟踪视觉注意力熵和输出图像相关性两个指标,根据异常情况触发对比对齐,实验证明该方法能提升模型可靠性,优于现有方法。

AI 中文摘要

人类视觉推理通常遵循从粗到细的注意力过程,从全局场景理解开始,逐渐聚焦于与问题相关的区域。然而,多模态大语言模型可能因注意力漂移和视觉证据利用不足而偏离此模式,导致幻觉。为缓解这些问题,本研究提出双指标引导对比对齐(DICA),在推理过程中跟踪两个信息论指标:视觉注意力熵(VAE),反映视觉注意力的集中程度;输出图像相关性(OIC),衡量生成输出对视觉输入的依赖性。VAE的异常增加或OIC的降低对应不同的失败模式,触发有针对性的对比对齐以恢复视觉基础。多个基准测试的实验结果表明,DICA始终优于现有方法,并大幅减少幻觉,突出了指标驱动干预在提高多模态推理可靠性方面的有效性。代码可在指定网址公开获取。

英文摘要

Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusing on question-relevant regions. However, multimodal large language models may deviate from this pattern due to attention drift and the underutilization of visual evidence, which can lead to hallucinations. To mitigate these issues, this study proposes a Dual-Indicator Guided Contrastive Alignment (DICA), which tracks two information-theoretic indicators during inference: Visual Attention Entropy (VAE), which reflects the concentration of visual attention, and Output Image Correlation (OIC), which measures the dependence of generated outputs on the visual input. An abnormal increase in VAE or a decrease in OIC corresponds to different failure modes, which trigger targeted contrastive alignment to restore visual grounding. Experimental results across multiple benchmarks demonstrate that DICA consistently outperforms existing approaches and substantially reduces hallucinations, highlighting the effectiveness of indicator-driven intervention in improving multimodal inference reliability. The code is publicly available at https://github.com/BGWH123/DICA/.

DOI:10.18653/v1/2026.findings-acl.1933

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑