arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05070cs.CV

MultiAttenGastro:用于胃肠内窥镜分类的多维注意力增强框架

MultiAttenGastro: Multi-Dimensional Attention Augmentation for Gastrointestinal Endoscopy Classification

Sadhana Devarajan, Praveen Kumar Chandaliya, Dhruvin Jashvant Kumar Shah, Kishor Upla, Kiran Raja

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出即插即用的多维注意力框架MultiAttenGastro,经跨数据集评估发现其效果与表征差距相关,仅在大差距场景下能提升多数骨干网络的胃肠内窥镜分类性能。

中文摘要 AI 辅助

自动胃肠(GI)内窥镜分类需要能够在不同模态和类别分布间泛化的模型,这类模型通常与自然图像预训练存在较大差异。我们提出MultiAttenGastro,这是一种即插即用的注意力框架,包含并行的1维通道、2维空间和3维上下文头;并在五个公开GI数据集上,对八个CNN和Transformer骨干网络开展了首次系统性跨数据集评估,共进行80次骨干-数据集运行。我们发现注意力的有效性并非通用,而是与ImageNet特征和目标分布之间的表征差距相关:MultiAttenGastro在Kvasir-Capsule(14类WCE,差距较大)上提升了8个骨干网络中的6个,最佳宏观F1值达98.33%;在差距较小的Kvasir-v2基准上,其效果均为负面(8个骨干网络提升数为0);在差距中等的数据集上则呈现混合结果。对最优情况(Kvasir-Capsule、ConvNeXt-Tiny)的五折消融实验表明,该提升在方向上一致,但统计上不具备决定性(配对t检验:p=0.47;Wilcoxon检验:p=0.63);且单个注意力头单独使用时并非均有益,仅组合使用才能产生正向平均效果。中心核对齐(CKA)分析将该模式与表征冗余度关联:在大领域差距下,头间CKA值低与框架仅有的一致增益重合;而在小差距下,高冗余度与框架的损失重合。我们报告这些结果(包括非显著边际),旨在说明多维注意力在胃肠内窥镜分类中何时及为何有效,而非声称MultiAttenGastro是绝对更优的架构选择。

英文摘要

Automated gastrointestinal (GI) endoscopy classification requires models that generalize across diverse modalities and class distributions, often far from natural-image pretraining. We propose MultiAttenGastro, a plug-and-play attention framework with parallel 1-D channel, 2-D spatial, and 3-D contextual heads, and present the first systematic cross-dataset evaluation across eight CNN and transformer backbones on five public GI datasets (80 backbone--dataset runs). We find that attention effectiveness is not universal but tracks the representational gap between ImageNet features and the target distribution: MultiAttenGastro improves 6 of 8 backbones on Kvasir-Capsule (14-class WCE, large gap; best macro F1 98.33\%), is uniformly negative on the small-gap Kvasir-v2 benchmark (0/8), and shows mixed outcomes on datasets with intermediate gap. Five-seed ablation on the strongest case (Kvasir-Capsule, ConvNeXt-Tiny) shows this improvement is directionally consistent, but not statistically decisive (paired $t$: $p=0.47$; Wilcoxon: $p=0.63$), and that individual attention heads are not uniformly beneficial in isolation only their combination yields a positive mean effect. Centered Kernel Alignment (CKA) analysis links this pattern to representational redundancy: low inter-head CKA under large domain gaps coincides with the framework's only consistent gains, while high redundancy under small gaps coincides with its losses. We report these results, including the non-significant margins, as evidence for when and why multi-dimensional attention helps GI endoscopy classification, rather than as a claim that MultiAttenGastro is a strictly superior architectural choice.

发表机构

  • Sardar Vallabhbhai National Institute of Technology(萨达尔·瓦拉巴伊·国家理工学院)
  • Swasthyam Gastro and Liver Hospital(斯瓦斯蒂亚姆胃肠肝病医院)
  • Norwegian University of Science and Technology(挪威科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑