发表机构
University of Warsaw; Centre for Credible AI, Warsaw University of Technology(华沙大学; 华沙理工大学可信人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视觉语言模型在医学任务中解释难的问题,引入ParseFIxLIP方法,将树图解析融入班扎夫交互博弈,通过smart_depth分组策略减轻概念碎片化,提升跨模态交互可解释性,为VLM决策提供医学领域相关见解。
AI 中文摘要
视觉语言模型(VLM)在医学任务中展现出强大能力,但在临床环境中进行可靠部署时,提供忠实且可解释的解释仍是关键考量。现有解释方法如FIxLIP框架,难以应对现代分词器的细粒度问题,导致临床概念碎片化,产生噪声和语义不连贯的跨模态归因。为此,我们引入ParseFIxLIP,将树图解析纳入FIxLIP使用的班扎夫交互博弈中。该语义感知策略利用依存句法分析树将相关文本令牌分组为语义连贯单元来定义解释参与者。我们的smart_depth分组策略根据spaCy令牌依存树合并令牌,成功减轻概念碎片化,通过统一复杂医学概念产生更具可解释性的跨模态交互。定量分析表明,我们的解析方法在长标题高维性问题上保持统计稳健性和语义简约性。对生物医学CLIP的定性分析在医学图像(ROCOv2)和一般示例上得到验证,证实该方法准确捕捉了分组单词对模型预测的协同影响。总之,我们的工作为VLM决策提供了直观且与临床相关的见解,满足了医学领域对连贯解释的迫切需求。
英文摘要
Vision-Language Models (VLMs) are demonstrating significant capabilities in medical tasks like radiology analysis, yet providing faithful and interpretable explanations remains a key consideration for their responsible deployment in clinical settings. However, existing explanation methods, such as the widely used FIxLIP framework, often struggle with the fine-grained nature of modern tokenizers. The tokenization problem fragments clinical concepts---splitting terms like "saddle embolus" into scattered, meaningless subwords---which leads to noisy, semantically incoherent cross-modal attributions. Such fragmentation also results in a combinatorial explosion of interaction possibilities, obscuring the model's true reasoning. To address this, we introduce ParseFIxLIP, an extension that incorporates the Tree-Gram Parsing into the Banzhaf interaction game used by FIxLIP. This semantically informed strategy utilizes dependency parsing trees to define explanation players by grouping related text tokens into semantically coherent units. Our smart_depth grouping strategy, merging tokens according to spaCy token dependency tree, successfully mitigates concept fragmentation, yielding substantially more interpretable cross-modal interactions by unifying complex medical concepts. Quantitatively, while baselines struggled with the high dimensionality of long captions, our parsing approach maintained statistical robustness and semantic parsimony. Qualitative analysis on BiomedCLIP, validated on medical imagery (ROCOv2) and general examples, confirms that the approach accurately captures the synergistic influence of grouped words on model predictions. In conclusion, our work offers intuitive and clinically relevant insights into VLM decision-making, fulfilling the critical need for coherent explanations in the medical domain.
Comments12 pages, 13 figures. Accepted at EXPLIMED 2026 (Third Workshop on Explainable Artificial Intelligence for the medical domain), IJCAI-ECAI 2026