发表机构
Eindhoven University of Technology; Amsterdam University Medical Centers; University of Amsterdam(埃因霍温理工大学; 阿姆斯特丹大学医学中心; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对内窥镜生成模型的领域差距与计算成本问题,提出REVEAL模型,基于500万帧内窥镜数据训练,在多任务中性能优于同类模型,为临床智能工具开发提供基础。
AI 中文摘要
开发内窥镜领域的基础生成模型面临两大限制:自然图像与临床图像之间的领域差距,以及训练大型扩散Transformer的计算成本。尽管表征对齐已提升了通用计算机视觉的效率,但其在高度专业化的内窥镜图像空间中的作用仍不明确。我们推出REVEAL(表征驱动型内窥镜视觉嵌入对齐),这是迄今为止最大的内窥镜专用生成基础模型,基于GastroNet-5M(GN-5M)训练,该数据集包含500万帧多中心内窥镜图像。REVEAL不依赖域外先验,而是采用直接在内窥镜分布上预训练的编码器,将扩散隐空间与领域特定视觉特征对齐,保留精细纹理与复杂解剖结构。除图像生成外,REVEAL还可作为强大的特征提取器;在多个基准测试中,其性能与专为分类任务微调的内窥镜基础模型(如EndoViT和Endo-FM)相当,部分场景下更优,且在真实成像干扰下展现出强大的表征鲁棒性。REVEAL生成高保真图像,且在隐空间编辑(如修复和扩展)中保持稳健的结构一致性。这一高容量骨干模型降低了构建专用临床工具的计算门槛,为未来智能胃肠病系统中的条件合成、分割和分布外检测提供了开放、通用的基础。
英文摘要
Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers. Although representation alignment has improved efficiency in general computer vision, its role within the highly specialized endoscopic image space remains unclear. We introduce REVEAL (Representation-driven Endoscopic Visual Embedding Alignment), the largest generative foundation model for endoscopy to date, trained on GastroNet-5M (GN-5M), a multicenter dataset of 5 million endoscopic frames. Instead of depending on out-of-domain priors, REVEAL employs encoders pretrained directly on the endoscopic distribution to align diffusion latents with domain-specific visual features, preserving fine textures and intricate anatomical structures. Beyond image generation, REVEAL also serves as a powerful feature extractor; in multiple benchmarks, it delivers performance that is competitive with, and in several cases exceeds, endoscopic foundation models such as EndoViT and Endo-FM, specifically tuned for classification tasks, while demonstrating strong representation robustness under realistic imaging corruptions. REVEAL produces high-fidelity images and maintains robust structural coherence in latent-space edits such as inpainting and outpainting. This high-capacity backbone lowers the computational threshold for building specialized clinical tools, offering an open, versatile foundation for conditional synthesis, segmentation, and out-of-distribution detection in future intelligent gastroenterology systems.
CommentsAccepted at the DCA-MI Workshop (ECCV 2026)