arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2508.07028cs.CV

基于大语言模型评估的独立注意力辅助图神经网络结合空间与结构信息交互的精确内窥镜图像分割

Large Language Model Evaluated Stand-alone Attention-Assisted Graph Neural Network with Spatial and Structural Information Interaction for Precise Endoscopic Image Segmentation

  • Department of Mathematics, The Chinese University of Hong Kong(香港中文大学数学系)
  • Department of Electronic and Computer Engineering, University of Massachusetts, Lowell(马萨诸塞大学低罗分校电子与计算机工程系)
  • Department of Radiology, Northwestern University(西北大学放射科)

机构由 AI 辅助整理,请以论文原文为准。

Juntong Fan, Shuyi Fan, Debesh Jha, Changsheng Fang, Tieyong Zeng, Hengyong Yu, Dayang Wang

更新

AI总结:

针对内窥镜息肉分割挑战,提出FOCUS-Med模型,通过双图卷积网络、独立自注意力和快速归一化融合策略整合空间与结构信息,并首次引入大语言模型进行定性评估,在五项关键指标上达到最先进水平。

AI中文摘要:

准确的息肉内窥镜图像分割对早期结直肠癌检测至关重要。然而,由于与周围黏膜对比度低、镜面高光和边界模糊,该任务仍具挑战性。为应对这些挑战,我们提出FOCUS-Med,代表内窥镜医学成像中空间与结构图融合及注意力上下文感知息肉分割。FOCUS-Med集成双图卷积网络(Dual-GCN)模块以捕获上下文空间和拓扑结构依赖关系。这种基于图的表示使模型能利用拓扑线索和空间连通性更好地区分息肉与背景组织,这些线索常在原始图像强度中被掩盖。它增强了模型保留边界和描绘息肉典型复杂形状的能力。此外,采用位置融合的独立自注意力来加强全局上下文整合。为弥合编码器-解码器层间的语义鸿沟,我们引入可训练的加权快速归一化融合策略以实现高效多尺度聚合。值得注意的是,我们首次引入大语言模型(LLM)对分割质量提供详细定性评估。在公开基准上的大量实验表明,FOCUS-Med在五项关键指标上达到最先进性能,凸显其在AI辅助结肠镜检查中的有效性和临床潜力。

英文摘要:

Accurate endoscopic image segmentation on the polyps is critical for early colorectal cancer detection. However, this task remains challenging due to low contrast with surrounding mucosa, specular highlights, and indistinct boundaries. To address these challenges, we propose FOCUS-Med, which stands for Fusion of spatial and structural graph with attentional context-aware polyp segmentation in endoscopic medical imaging. FOCUS-Med integrates a Dual Graph Convolutional Network (Dual-GCN) module to capture contextual spatial and topological structural dependencies. This graph-based representation enables the model to better distinguish polyps from background tissues by leveraging topological cues and spatial connectivity, which are often obscured in raw image intensities. It enhances the model's ability to preserve boundaries and delineate complex shapes typical of polyps. In addition, a location-fused stand-alone self-attention is employed to strengthen global context integration. To bridge the semantic gap between encoder-decoder layers, we incorporate a trainable weighted fast normalized fusion strategy for efficient multi-scale aggregation. Notably, we are the first to introduce the use of a Large Language Model (LLM) to provide detailed qualitative evaluations of segmentation quality. Extensive experiments on public benchmarks demonstrate that FOCUS-Med achieves state-of-the-art performance across five key metrics, underscoring its effectiveness and clinical potential for AI-assisted colonoscopy.

补充信息

↑