arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24730cs.CVcs.AI

KANEx:将柯尔莫哥洛夫-阿诺德网络的可解释性转化为医学可解释性

KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

  • Indian Institute of Technology Madras(印度理工学院马德拉斯分校)
  • Vanderbilt University School of Medicine(范德堡大学医学院)

机构由 AI 辅助整理,请以论文原文为准。

Krithi Shailya, Ananya Lakshmi Ravi, Venkatanathan K. V., Sowmya S. Sundaram, Gokul S. Krishnan, Aditi Anand, Balaraman Ravindran

AI总结:

研究针对医学应用中视觉模型黑箱问题,提出KANEx框架,利用柯尔莫哥洛夫-阿诺德网络的符号透明度为VLM推理奠基,设计KAN-Map热图生成方法,经实验验证该方法能提升语义相似度、视觉定位及推理质量,为可信医学人工智能发展助力。

AI中文摘要:

计算机视觉模型在医学应用中已非常有效,但其黑箱性质削弱了临床医生的信任。在临床工作流程中,胸部X光分类器 increasingly 与视觉语言模型(VLM)结合以生成自然语言解释。然而,这些系统增加了语言流畅性,却未解决视觉模型的潜在不透明性。随着基于样条的组件提供固有可解释功能单元的柯尔莫哥洛夫-阿诺德网络(KAN)的出现,研究能否利用其架构透明度来产生更可信的文本解释。介绍了KANEx,首个利用KAN的符号透明度为VLM推理奠定基础的框架。还设计了KAN-Map,一种直接从KAN模型而非梯度近似得出的新型热图生成方法。将这些有基础的上下文输入下游VLM以增强可解释性。在MIMIC-CXR数据集上进行基准测试,结果表明基于KAN的架构与ResNet/ViT基线相比,语义相似度提高,同时生成更忠实的显著性图,视觉定位和下游推理质量提高了10%。研究结果表明,将语言解释和视觉归因建立在数学可解释单元上是迈向可信医学人工智能的必要一步。

英文摘要:

Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In clinical workflows, chest X-ray classifiers are increasingly paired with Vision-Language Models (VLMs) to generate natural-language explanations. However, these systems add linguistic fluency without addressing the underlying opacity of the visual model. With the emergence of Kolmogorov-Arnold Networks (KANs), whose spline-based components provide inherently interpretable functional units, we investigate whether this architectural transparency can be leveraged to produce more trustworthy textual explanations. We introduce KANEx, the first ever framework that leverages the symbolic transparency of KANs to ground VLM reasoning. This interpretability also made it possible to design KAN-Map, a novel heatmap generation method derived directly from KAN models rather than gradient approximations. We feed these grounded contexts into downstream VLMs for enhanced explainability. Benchmarked on the MIMIC-CXR dataset, we demonstrate that KAN-based architectures with ResNet/ViT baselines demonstrate improved semantic similarity while producing significantly more faithful saliency maps. KAN architectures improve visual localization and downstream reasoning quality by 10%. Our findings suggest that grounding linguistic explanations and visual attributions in mathematically interpretable units is a necessary step toward trustworthy medical AI.

补充信息

↑