CEM-TUDASR:基于Transformer的计算高效多模态无监督域自适应超分辨率方法
CEM-TUDASR: Computationally efficient multi-modality transformer based unsupervised domain adaptive super-resolution approach
浏览论文内容
中文总结 AI 辅助
针对无线胶囊内窥镜低分辨率图像,提出基于Transformer的无监督域自适应超分辨率框架CEM-TUDASR,通过域自适应退化网络和注意力模块,实现高效高质量重建,参数少且计算量低。
中文摘要 AI 辅助
无线胶囊内窥镜(WCE)能够实现胃肠道的无创可视化,但其微型化光学元件、传感器限制以及无线传输约束导致图像分辨率低,诊断重要结构的可见性降低。本文提出CEM-TUDASR,一种计算高效的无监督基于Transformer的超分辨率框架,用于WCE图像增强,无需配对的低分辨率(LR)和高分辨率(HR)训练数据。域自适应退化网络从HR常规内窥镜图像合成逼真的WCE样LR图像,缩小域差距,实现有效的非配对学习。SR生成器集成深度注意力块(DABs)和融合注意力块(FAB),以捕获长距离上下文依赖和精细局部结构,同时保持感知和结构保真度。模型在源自Kvasir Capsule的精选数据集上训练,并在KID和GIANA上评估跨数据集泛化能力。无参考质量指标,包括BRISQUE、PIQE、NIQE和特定领域的EndoQM,表明CEM-TUDASR始终优于现有无监督SR方法。定性结果进一步证明黏膜纹理、血管模式和临床相关解剖细节的恢复得到改善。视网膜图像的跨域实验进一步证明了该框架的适应性。仅2.67百万参数和169.94 GFLOPs,CEM-TUDASR在保持计算效率的同时实现高质量重建,适用于资源受限的临床和嵌入式内窥镜应用。
英文摘要
Wireless Capsule Endoscopy (WCE) enables non-invasive visualization of the gastrointestinal tract, but its miniaturized optics, sensor limitations, and wireless transmission constraints result in low-resolution images with reduced visibility of diagnostically important structures. This paper proposes CEM-TUDASR, a computationally efficient unsupervised Transformer-based super-resolution framework for WCE image enhancement without paired low-resolution (LR) and high-resolution (HR) training data. A domain-adaptive degradation network synthesizes realistic WCE-like LR images from HR conventional endoscopy images, reducing the domain gap and enabling effective unpaired learning. The SR generator integrates Deep Attention Blocks (DABs) and a Fusion Attention Block (FAB) to capture long-range contextual dependencies and fine local structures while preserving perceptual and structural fidelity. The model is trained on a curated dataset derived from Kvasir Capsule and evaluated on KID and GIANA for cross-dataset generalization. No-reference quality metrics, including BRISQUE, PIQE, NIQE, and the domain-specific EndoQM, show that CEM-TUDASR consistently outperforms existing unsupervised SR methods. Qualitative results further demonstrate improved restoration of mucosal textures, vascular patterns, and clinically relevant anatomical details. Cross-domain experiments on retinal images additionally demonstrate the adaptability of the framework. With only 2.67 million parameters and 169.94 GFLOPs, CEM-TUDASR achieves high-quality reconstruction while maintaining computational efficiency, making it suitable for resource-constrained clinical and embedded endoscopic applications.
发表机构
- Sardar Vallabhbhai National Institute of Technology (SVNIT)(萨达尔·瓦拉巴伊国家技术学院)
- Norwegian University of Science and Technology (NTNU)(挪威科技大学)
机构由 AI 辅助整理,请以论文原文为准。