MVC:一种用于胸部X线影像COVID-19诊断的多任务视觉Transformer网络
MVC: A Multi-Task Vision Transformer Network for COVID-19 Diagnosis from Chest X-ray Images
- Deakin University(迪肯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有医学图像诊断技术多为单任务、缺乏统一多任务框架的问题,研究人员提出多任务视觉Transformer方法MVC,可同时完成胸部X光片分类与感染区域识别,在COVID-19基准数据集上表现优于现有基线。
AI中文摘要:
过去十年间,利用计算机算法开展医学图像分析受到了研究界的广泛关注,并取得了巨大进展。随着计算资源的最新发展以及大规模医学图像数据集的可获得性提升,研究人员已开发出众多深度学习模型,用于从医学图像中进行疾病诊断。然而,现有技术分别聚焦于疾病分类、病灶识别等子任务,缺乏可实现多任务诊断的统一框架。受视觉Transformer在局部与全局表征学习方面的能力启发,本文提出了一种名为Multi-task Vision Transformer(MVC,多任务视觉Transformer)的新方法,可同时对胸部X线图像进行分类,并从输入数据中识别受感染区域。该方法以视觉Transformer为基础,在多任务场景下拓展了其学习能力。我们在一个COVID-19胸部X线图像基准数据集上对所提方法进行了评估,并与现有基线模型展开对比。实验结果证实,所提方法在图像分类和受感染区域识别两项任务上均优于基线模型。
英文摘要:
Medical image analysis using computer-based algorithms has attracted considerable attention from the research community and achieved tremendous progress in the last decade. With recent advances in computing resources and availability of large-scale medical image datasets, many deep learning models have been developed for disease diagnosis from medical images. However, existing techniques focus on sub-tasks, e.g., disease classification and identification, individually, while there is a lack of a unified framework enabling multi-task diagnosis. Inspired by the capability of Vision Transformers in both local and global representation learning, we propose in this paper a new method, namely Multi-task Vision Transformer (MVC) for simultaneously classifying chest X-ray images and identifying affected regions from the input data. Our method is built upon the Vision Transformer but extends its learning capability in a multi-task setting. We evaluated our proposed method and compared it with existing baselines on a benchmark dataset of COVID-19 chest X-ray images. Experimental results verified the superiority of the proposed method over the baselines on both the image classification and affected region identification tasks.