AI 中文总结
提出基于原型的模糊规则框架,无需微调即可解释胃肠基础模型的内部特征,在多个任务上达到与黑盒分类器相当的准确率,并用于分析合成图像。
AI 中文摘要
在大规模数据集上预训练的基础模型展示了强大的医学影像任务迁移能力。然而,理解其潜在表征如何编码临床相关信息,在安全关键领域仍是一个开放的挑战。本研究提出了一种基于原型(prototype)的模糊规则框架,无需任何微调,即可解释预训练基础模型内部层产生的补丁级特征。通过在特征空间中进行聚类学习类别特定的原型,生成紧凑的可视模式。补丁特征随后被表示为原型相似度,并通过具有人类可读的IF-THEN语言条件的模糊规则进行分类。该框架应用于在ImageNet-1K和GastroNet-5M上预训练的ViT-S/16骨干网络的最后两个块,并在相同的冻结特征下,与k近邻、核SVM和线性探针进行基准比较,涉及无线胶囊内窥镜分类、胃肠内窥镜分类和结肠息肉分割。实验分析表明,所提方法在无需骨干微调的情况下,达到了与这些黑盒分类器相当的准确率,并且领域特定的预训练产生了既具有判别性又具有符号可压缩性的特征。由于所得规则是从真实数据中提取并以可解释的术语表达,它们进一步被用作研究合成医学图像的工具,提供了人类可读的说明,展示生成器复现或未能复现哪些真实原型和规则,定位合成图像偏离真实组织的位置,而不是用单一分数进行概括。该框架因此提供了基础模型如何组织临床相关结构的透明、深度分辨视图,以及提取规则的实际下游用途。
英文摘要
Foundation models pretrained on large-scale datasets demonstrate strong transferability to medical imaging tasks. However, understanding how their latent representations encode clinically relevant information remains an open challenge in safety-critical domains. This study proposes a prototype-based fuzzy-rule framework that interprets the patch-level features produced by the inner layers of pretrained foundation models, without any fine-tuning. Class-specific prototypes are learned by clustering in the feature space, yielding compact visual patterns. Patch features are then expressed as prototype similarities and classified by fuzzy rules with linguistic IF-THEN conditions that are human readable. The framework is applied across the final two blocks of ViT-S/16 backbones pretrained on ImageNet-1K and GastroNet-5M, and benchmarked against k-nearest neighbours, kernel SVM, and linear probing under identical frozen features, on wireless capsule endoscopy classification, gastrointestinal endoscopy classification, and colonic polyp segmentation. The experimental analysis shows that the proposed method, without backbone fine-tuning, reaches accuracy comparable to these black-box classifiers, and that domain-specific pretraining yields features that are both discriminative and symbolically compressible. Because the resulting rules are extracted from real data and expressed in interpretable terms, they are further used as an instrument to investigate synthetic medical images, providing a human-readable account of which real prototypes and rules a generator reproduces or fails to reproduce, localising where a synthetic image departs from real tissue rather than summarising it with a single score. The framework thus offers a transparent, depth-resolved view of how foundation models organise clinically relevant structure, together with a practical downstream use of the extracted rules.
CommentsAccepted at the excv, ECCV 2026 Workshops. 17 pages, 4 figures, 7 tables