arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CRAFT:医学视觉语言模型中的因果责任与失败追踪

CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models

Chunzheng Zhu, Jiaqi Zeng, Hongbo Zhao, Yihang Chen, Yijun Wang, Jianxin Lin

arXiv 2609.38810首次发表:更新:

发表机构

Hunan University; University of Hong Kong(湖南大学; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对医学视觉语言模型中的仲裁失败和刹车失败,提出CRAFT方法,通过因果头定位与干预,实现精准识别和纠正,提升诊断安全性。

AI 中文摘要

随着视觉语言模型越来越多地部署于临床诊断,理解其内部如何解决相互竞争的视觉和文本信号成为一项安全要务。现有的机制分析仍局限于单模态文本,且无法解释为何一句误导性句子能推翻基于正确图像的诊断,或为何模型在视觉证据不足时仍给出自信答案。我们发现这两种安全风险——即文本上下文覆盖视觉基础的仲裁失败,以及模型在证据不足时仍做出承诺的刹车失败——由空间上不相交的注意力头群体介导:仲裁头形成一条中到深的宽带,反映跨层证据竞争;而刹车头则集中在较窄的中后层带,调节证据充分性和弃权(不执行)行为。为将这些观察建立在因果电路上,我们引入了CRAFT,通过双重标准将每种失败模式定位到最小的因果头集,并通过时间探针和Tuned Lens轨迹分析验证其必要性和充分性。切除仲裁头可显著减少冲突跟随,而对干净输入的退化可忽略不计;切除刹车头则能在视觉证据退化时恢复适当的弃权(不执行)行为。这两种干预针对空间上不相交的头集,并产生不同的纠正效果,强调了失败模式的机制可分离性。在多个医学VQA基准和VLM架构上的实验验证了定位和干预的有效性,表明所识别的头在因果上驱动每种失败模式,且有针对性的调制无需重新训练即可泛化。代码可在GitHub仓库获取。

英文摘要

As vision language models are increasingly deployed in clinical diagnosis, under standing how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitra tion failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores ap propriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and inter ventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at https://github.com/zhcz328/CRAFT.

CommentsNeurIPS 2026 Spotlight, Medical VLM Failure Analysis

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑