发表机构
IEM Saltlake; Techno Main Salt Lake; VIT Bhopal University; IIT Jodhpur(IEM盐湖校区; Techno主盐湖校区; 博帕尔维洛尔理工大学; 焦特布尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出OptiModNet,一款用于视盘和视杯分割的轻量混合架构,结合分组查询与通道注意力及聚合金字塔损失,在REFUGE2数据集上实现超现有方法2.5%的最优性能,且仅需3.73 GFLOPs和1.93M参数,兼顾高性能与低计算量。
AI 中文摘要
视盘和视杯的精确分割对于青光眼的早期检测与诊断至关重要。然而,在跨数据集保持一致的高性能表现的同时,维持较低的计算需求仍是一项重大挑战。在青光眼检测中,低计算量方法对于实现快速大规模筛查、推动其在资源受限的临床环境中部署至关重要。尽管UNet、视觉Transformer(ViT)、扩散模型等深度学习模型已展现出强大的分割性能,但这些方法通常会带来巨大的计算开销:UNet擅长捕捉局部特征,但在建模全局上下文信息方面存在局限;相反,ViT擅长长程依赖建模,但计算密集度高。混合架构如UNetR,将基于Transformer的编码器与UNet风格的解码器相结合,已展现出性能提升,但会增加额外的复杂度。鉴于此,本研究提出OptiModNet,一款专为视盘和视杯分割设计的轻量新型混合架构。该模型在网络的多个阶段整合多种注意力机制,以增强局部和全局特征表示;同时引入聚合金字塔损失,在解码器的多个深度层级对预测结果进行监督,以促进更好的梯度流动和结构一致性。我们在REFUGE2数据集上对OptiModNet的视盘和视杯分割任务进行评估,结果显示,所提方法达到了当前最优性能,较现有方法提升超过2.5%,同时保持了高计算效率,仅需3.73 GFLOPs,参数规模为1.93M。代码可在该URL获取。
英文摘要
Precise segmentation of the optic disc and cup is critical for the early detection and diagnosis of glaucoma. However, achieving consistently high performance across datasets while maintaining low computational requirements remains a significant challenge. In glaucoma detection, low-computation methods are crucial for enabling rapid, large-scale screening and facilitating deployment in resource-limited clinical environments. While deep learning models such as UNets, Vision Transformers (ViTs), and Diffusion models have demonstrated strong segmentation performance but these methods often come with substantial computational overhead. UNets are efficient at capturing local features but are limited in modeling global contextual information. Conversely, ViTs excel at long-range dependency modeling but are computationally intensive. Hybrid architectures, such as UNetR, which combine transformer-based encoders with UNet-style decoders, have shown improved performance but while incurring additional complexity. Considering these, in this work, we propose OptiModNet, a light weight novel hybrid architecture tailored for optic disc and cup segmentation. The model integrates diverse attention mechanisms at multiple stages of the network to enhance both local and global feature representation. We include an Aggregated Pyramid Loss that supervises predictions at multiple decoder depths, to promote better gradient flow and structural consistency. We evaluate OptiModNet on the REFUGE2 dataset for both optic disc and cup segmentation tasks. Our method achieves state-of-the-art performance, exceeding existing approaches by over 2.5\%, while maintaining high efficiency with only 3.73 GFLOPs and 1.93M parameters. The code is available at https://github.com/SG1947/OptiModNet.