arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23853cs.CV

LUX:一种用于可解释内镜图像描述的病灶感知图条件视觉-语言架构

LUX: A Lesion-Aware Graph-Conditioned Visual - Language Architecture for Explainable Endoscopic Captioning

  • School of Engineering and Sciences, Tecnologico de Monterrey(蒙特雷理工学院工程与科学学院)
  • School of Computer Science, University of Leeds(利兹大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

Alexis Ivan Escamilla-Lopez, Gilberto Ochoa-Ruiz, Salvador Hinojosa, Sharib Ali

AI总结:

该研究针对内镜图像解读的主观性问题,提出LUX图条件视觉-语言架构,通过病灶中心场景图实现可解释描述,在多指标上优于现有模型并减少了幻觉与定位误差。

AI中文摘要:

溃疡性结肠炎的内镜图像解读复杂且主观,存在人为评估差异和黏膜炎症的细微变化。尽管深度学习推动了自动化分析的进展,但多数视觉-语言模型依赖全局视觉嵌入,忽略了病理证据的局部性和关联性,限制了临床可靠性与可解释性。本文提出LUX(病灶感知统一可解释描述模型,Lesion-aware Unified eXplainable captioning),一种用于可解释内镜图像描述的图条件视觉-语言架构。LUX基于Grad-CAM和CBAM激活图构建以病灶为中心的场景图,将病理区域表示为节点并编码其空间与临床关系。这些图嵌入被整合到T5解码器的交叉注意力层,使生成的单词能关注特定病灶节点而非仅全局图像特征,实现语言内容与病理证据的直接对齐,支持令牌级可解释性和关系推理。LUX在BLEU、METEOR、ROUGE-L和CIDEr指标上均优于强基线及现有最优医学描述模型,尤其在CIDEr上提升显著;同时减少了幻觉式临床发现,通过生成令牌与局部病理区域的更强对应关系改善了病灶级定位效果。

英文摘要:

The interpretation of endoscopic imagery in ulcerative colitis is complex and subjective, with variability in human assessment and subtle mucosal inflammation. Although deep learning has advanced automated analysis, most vision-language models rely on global visual embeddings that overlook the localized and relational nature of pathological evidence, limiting clinical reliability and interpretability. We introduce LUX (Lesion-aware Unified eXplainable captioning), a graph-conditioned vision-language architecture for explainable endoscopic image captioning. LUX constructs a lesion-centric scene graph from Grad-CAM and CBAM activation maps, representing pathological regions as nodes and encoding their spatial and clinical relationships. These graph embeddings are integrated into the cross-attention layers of a T5 decoder, enabling generated words to attend to specific lesion nodes rather than only to global image features. This provides direct alignment between linguistic content and pathological evidence, supporting token-level interpretability and relational reasoning. LUX outperforms strong baseline and state-of-the-art medical captioning models across BLEU, METEOR, ROUGE-L, and CIDEr, with particularly strong gains in CIDEr. It also reduces hallucinated clinical findings and improves lesion-level grounding through stronger correspondence between generated tokens and localized pathological regions.

↑