arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正确答案,错误依据:面向红外图像多模态大语言模型的感知感知评估与热基础反馈机制

Right Answer, Wrong Heat: Explanation-Aware Evaluation and Thermal-Grounded Feedback for MLLMs on Infrared Images

Yongsong Huang, Xiaofeng Liu, Tomo Miyazaki, Yaohou Fan, Shinichiro Omachi

arXiv 2608.09145首次发表:更新:

AI 中文总结

该研究针对红外图像多模态大语言模型仅以答案准确率评估的问题,提出感知感知评估框架与无需训练的热基础反馈机制,可提升解释热基础度,为可信红外场景理解模型的开发提供了方向。

AI 中文摘要

通用多模态大语言模型(MLLMs)正越来越多地应用于红外图像领域,目前通常仅以答案准确率对其进行评估。然而,正确的答案并不意味着模型的解释是基于红外热证据的。我们提出了一种感知感知评估框架,该框架将答案正确性、输出级解释的基础度以及红外视觉问题的热基础度分离开来。通过采用双LLM共识评判器并结合初步的人工锚定校准检查,我们发现:正确答案仍可能依赖于弱证据或可见光证据;移除原始红外图像并仅显示类可见光渲染会削弱热基础度,但准确率几乎没有变化;这种削弱在能力更强的模型中表现最为明显,但在保留红外图像时则会消失。我们进一步提出了热基础反馈(Thermal-Grounded Feedback,TGF),这是一种无需训练的反馈循环,可诊断解释层面的故障并在保留所选答案的同时修正解释。在本地配对输入验证中,TGF提升了解释层面的基础度且未改变答案。这些发现表明,未来用于红外场景理解的可信MLLMs应接受评估并被开发为能够生成基于热的解释,而非仅仅是准确的答案。

英文摘要

General-purpose multimodal large language models (MLLMs) are increasingly applied to infrared images, where they are commonly scored by answer accuracy alone. However, a correct answer does not ensure that the model's explanation is grounded in infrared thermal evidence. We introduce an explanation-aware evaluation framework that separates answer correctness, output-level explanation groundedness, and thermal grounding for infrared visual questions. Using a Dual-LLM Consensus Judge with a preliminary human-anchor calibration check, we find that correct answers can still rely on weak or visible-light evidence; withholding the original infrared image and showing only a visible-like rendering erodes thermal grounding with little accuracy change; and this erosion is observed most strongly for more capable models but disappears when infrared remains available. We further propose Thermal-Grounded Feedback (TGF), a training-free feedback loop that diagnoses explanation-side failures and revises the explanation while preserving the selected answer. On local paired-input validation, TGF improves explanation-side grounding without changing answers. These findings suggest that future trustworthy MLLMs for infrared scene understanding should be evaluated and developed to produce thermally grounded explanations rather than merely accurate answers.

CommentsThis manuscript is currently under peer review. Copyright may subsequently be transferred to the publisher, after which the availability of this version may be subject to the publisher's policy

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑