Object detection and segmentation in three-dimensional medical images is a very active area of research. However, most proposed deep learning models carry a high computational cost, and only few aim to be broadly applicable, achieve high detection performance, and remain fast to execute on resource-constrained hardware. To address this gap, we present RadYOLO, a 3D extension of YOLO11 tailored to medical images. We compare it with nnU-Net and nnDetection on five datasets comprising CT and MRI data with varying object sizes and prevalence. RadYOLO's detection performance surpasses that of nnDetection on four of five datasets and is comparable on one. Compared to nnU-Net, RadYOLO performs better on lesion detection tasks, while nnU-Net excels at detecting large organs when precise localization is required. When rough object localization is sufficient, RadYOLO matches or outperforms nnU-Net on all five datasets. Regarding inference time, RadYOLO is 8-46x faster than nnU-Net on a GPU. Compared to nnDetection the speedup is even higher. When executed on a CPU, RadYOLO's inference runs within seconds (still faster than nnU-Net on a GPU) offering a significant advantage for clinical and edge-device deployment.
RadYOLO repository: https://github.com/FraunhoferMEVIS/RadYOLO
CommentsThis preprint has been accepted for publication in the proceedings of the IEEE 23rd International Symposium on Biomedical Imaging (ISBI 2026). The final published version is available at https://doi.org/10.1109/ISBI61048.2026.11515734. The copyright for this work has been transferred to IEEE
Journal refProceedings of the 23rd IEEE International Symposium on Biomedical Imaging (ISBI), 2026
MRI provides superior soft tissue contrast without ionizing radiation; however, the absence of electron density information limits its direct use for dose calculation. As a result, current radiotherapy workflows rely on combined MRI and CT acquisitions, increasing registration uncertainty and procedural complexity. Synthetic CT generation enables MRI only planning but remains challenging due to nonlinear MRI-CT relationships and anatomical variability. We propose Parallel Swin Transformer-Enhanced Med2Transformer, a 3D architecture that integrates convolutional encoding with dual Swin Transformer branches to model both local anatomical detail and long-range contextual dependencies. Multi-scale shifted window attention with hierarchical feature aggregation improves anatomical fidelity. Experiments on public and clinical datasets demonstrate higher image similarity and improved geometric accuracy compared with baseline methods. Dosimetric evaluation shows clinically acceptable performance, with a mean target dose error of 1.69%. Code is available at: https://github.com/mobaidoctor/med2transformer.
Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research
通过用于MRI到CT转换的自适应编码器冻结实现节能联邦学习:一项绿色人工智能引导的研究
Ciro Benito Raggio, Lucia Migliorelli, Nils Skupien, Mathias Krohmer Zabaleta, Oliver Blanck, Francesco Cicone, Giuseppe Lucio Cascini, Paolo Zaffino, Maria Francesca Spadea
机构
*
Institute of Biomedical Engineering(生物医学工程研究所)
;
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
Department of Political Science(政治学系)
;
Università Degli Studi Di Teramo(特拉莫大学)
;
Department of Radiation Oncology(放射肿瘤学系)
;
University Medical Center Schleswig-Holstein(什未林施瓦茨-霍斯特大学医院)
;
Department of Experimental and Clinical Medicine(实验与临床医学系)
;
Magna Graecia University(马格拉西亚大学)
Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions to collaboratively train deep learning (DL) models, even with limited data. However, the significant resource requirements of FL often exclude centres with limited computational infrastructure, further widening existing healthcare disparities. To address this issue, we propose a Green AI-oriented adaptive layer-freezing strategy designed to reduce energy consumption and computational load while maintaining model performance. We tested our approach using different federated architectures for Magnetic Resonance Imaging (MRI)-to-Computed Tomography (CT) conversion. The proposed adaptive strategy optimises the federated training by selectively freezing the encoder weights based on the monitored relative difference of the encoder weights from round to round. A patience-based mechanism ensures that freezing only occurs when updates remain consistently minimal. The energy consumption and CO2eq emissions of the federation were tracked using the CodeCarbon library. Compared to equivalent non-frozen counterparts, our approach reduced training time, total energy consumption and CO2eq emissions by up to 23%. At the same time, the MRI-to-CT conversion performance was maintained, with only small variations in the Mean Absolute Error (MAE). Notably, for three out of the five evaluated architectures, no statistically significant differences were observed, while two architectures exhibited statistically significant improvements. Our work aligns with a research paradigm that promotes DL-based frameworks meeting clinical requirements while ensuring climatic, social, and economic sustainability. It lays the groundwork for novel FL evaluation frameworks, advancing privacy, equity and, more broadly, justice in AI-driven healthcare.
A Proof-of-Concept Study of Multitask Learning for Cranial Synthetic CT Generation Across Heterogeneous MRI Field Strengths
多任务学习在跨异质MRI场强的颅骨合成CT生成中的概念验证研究
Zhuoyao Xin, Yiren Zhang, Christopher Wu, Dong Liu, Chunming Gu, Elena Greco, Erik H. Middlebrooks, Jun Hua, Jia Guo
机构
*
F.M. Kirby Research Center for Brain Imaging, Kennedy Krieger Institute(F.M. Kirby脑成像研究中心,Kennedy Krieger研究所)
;
Neurosection, Division of MR Research, Russell H. Morgan Department of Radiology and Radiological Science, Johns Hopkins University School of Medicine(神经科,磁共振研究部,约翰·霍普金斯大学医学院放射学与放射科学系)
;
Department of Biomedical Engineering, Johns Hopkins University(生物医学工程系,约翰·霍普金斯大学)
;
Department of Biomedical Engineering, Case Western Reserve University(生物医学工程系,凯斯西储大学)
;
Department of Biomedical Engineering, Columbia University(生物医学工程系,哥伦比亚大学)
;
Department of Neuroscience, Columbia University(神经科学系,哥伦比亚大学)
;
Department of Radiology, Mayo Clinic(放射科,梅奥诊所)
;
Neuroradiology and Neurosurgery, Mayo Clinic College of Medicine and Science(神经放射学与神经外科,梅奥诊所医学院与科学学院)
Accurate synthesis of computed tomography (CT) images from magnetic resonance imaging (MRI) is clinically valuable for cranial applications such as attenuation correction, radiotherapy planning, and image-guided interventions. However, heterogeneity across MRI field strengths and acquisition protocols limits the generalizability of existing methods. In this study, we formulate cranial CT synthesis as a modular, structurally coupled problem and propose a deep learning framework to improve robustness across heterogeneous MRI conditions. The model is designed to adapt to variations in field strength and imaging protocols while preserving anatomical consistency. Experiments on multi-site datasets demonstrate improved performance and generalization compared with conventional approaches. The proposed method enables reliable CT synthesis across heterogeneous MRI settings, supporting broader clinical translation.
MSR:Hybrid Field Modeling for CT-MRI Rigid-Deformable Registration of the Cervical Spine with an Annotated Dataset
MSR:用于带有标注数据集的颈椎脊柱CT-MRI刚-变形配准的混合场建模
Bohai Zhang, Wenjie Chen, Mu Li, Kaixing Long, Xing Shen, Xinqiang Yao, Jincheng Yang, Jianting Chen, Wei Yang, Qianjin Feng, Lei Cao
机构
*
School of Biomedical Engineering, Southern Medical University, Guangzhou, 510515, China.(南方医科大学生物医学工程学院)
;
Guangdong Provincial Key Laboratory of Medical Image Processing, Guangzhou, 510515, China.(广东省医学图像处理重点实验室)
;
Guangdong Province Engineering Laboratory for Medical Imaging(广东省医学影像与诊断技术工程实验室)
;
Information Center, Nanfang Hospital, Southern Medical University, Guangzhou, 510515, China.(南方医科大学南华医院信息中心)
;
Division of Spine Surgery, Department of Orthopaedics, Nanfang hospital, Southern Medical University, Guangzhou, Guangdong, 510515, China.(南方医科大学南华医院骨科脊柱外科)
Accurate CT-MRI registration of the cervical spine is essential for preoperative planning because this region is anatomically complex,highly variable,and vulnerable to injury of the vertebral arteries and spinal cord. However,cervical CT-MRI registration remains underexplored,particularly for rigid-deformable hybrid modeling,and the lack of high-quality annotated multimodal data further limits progress. To address these challenges, we construct and release a comprehensively annotated CT-MRI dataset, R-D-Reg, and propose MSR, a rigid-deformable hybrid registration framework for complex joint structures. Specifically, MSR includes a rigid registration module for independent local rigid alignment of individual vertebrae and a deformable registration module with an MSL block that combines Mamba-based global modeling and Swin Transformer-based local modeling through adaptive gating. The rigid and deformable deformation fields are then fused to generate a hybrid field that better preserves local anatomical consistency. The code and dataset are publicly available at https://github.com/ssc1230609-spec/MSR-registration.
Generating synthetic CT (sCT) from MRI or CBCT plays a crucial role in enabling MRI-only and CBCT-based adaptive radiotherapy, improving treatment precision while reducing patient radiation exposure. To address this task, we adopt a fully 3D Flow Matching (FM) framework, motivated by recent work demonstrating FM's efficiency in producing high-quality images. In our approach, a Gaussian noise volume is transformed into an sCT image by integrating a learned FM velocity field, conditioned on features extracted from the input MRI or CBCT using a lightweight 3D encoder. We evaluated the method on the SynthRAD2025 Challenge benchmark, training separate models for MRI to sCT and CBCT to sCT across three anatomical regions: abdomen, head and neck, and thorax. Validation and testing were performed through the challenge submission system. The results indicate that the method accurately reconstructs global anatomical structures; however, preservation of fine details was limited, primarily due to the relatively low training resolution imposed by memory and runtime constraints. Future work will explore patch-based training and latent-space flow models to improve resolution and local structural fidelity.
Objective: To determine whether free-breathing golden-angle radial sparse parallel (GRASP) magnetic resonance imaging (MRI) can represent respiratory-induced organ motion in patients with liver malignancies undergoing stereotactic body radiation therapy (SBRT). Methods: A retrospective analysis of 54 patients undergoing liver SBRT was conducted. Four-dimensional computed tomography (4D-CT), the gold standard for motion assessment, was used to characterize liver tumor motion. Image fusion was performed between free-breathing GRASP MRI and each respiratory phase of 4D-CT using an in-house registration program, with fusion quality quantified by maximum cross-correlation coefficient (MCC). Validation involved two blinded radiation oncologists: one repeated image fusion using the Eclipse-built-in module, while the other evaluated clinical relevance on a five-point scale. Results: The 50% respiratory phase of 4D-CT achieved the highest fusion quality with GRASP MRI, showing no significant differences compared to the 30% (P = 0.106), 40% (P = 0.632), and 60% (P = 0.792) phases. In contrast, fusion quality declined significantly beyond the mid-respiratory window (30%-60%), with poor fusion at the 0%, 10%, 20%, 80%, and 90% phases (P < 0.001). Validation by radiation oncologists corroborated these findings, with the 50% phase achieving the highest score. Subjective scores remained above 4 for phases 30%-70%, while scores for the remaining phases fell below 4. Conclusion: Free-breathing GRASP MRI cannot independently represent organ motion across all respiratory phases; it accurately characterizes motion only within the mid-respiratory phases (30%-60%), with optimal performance at the 50% phase. When used as a delineation standard in liver SBRT, GRASP MRI should be combined with 4D-CT or dynamic imaging modalities to ensure comprehensive motion assessment and accurate target volume definition.
Automated diagnosis of 3D brain CT scans is essential for critical care, yet it remains challenging due to the heavy reliance on manual annotations and the limited semantic understanding of conventional models. While 2D foundation vision-language models (VLMs) have shown remarkable generalization, effectively transferring their representational power to 3D volumes remains an open problem. In this paper, we propose Brain-Adapter, a novel dual-stream multiple instance learning (MIL) framework that leverages pre-trained 2D biomedical VLMs and raw diagnostic reports for robust scan-level multi-label classification. Specifically, we introduce a Text-Conditioned Attention (TCA) mechanism, utilizing raw diagnostic sentences as semantic queries to dynamically align visual cues with specific disease concepts. Concurrently, a parallel visual MIL stream captures global scan characteristics, supervised by structured labels extracted via a Large Language Model (LLM). To ensure representation coherence, a consistency constraint enforces synergy between the two streams. During inference, an Uncertainty-Aware Refinement (UAR) module dynamically calibrates and fuses these dual-stream predictions to resolve ambiguous cases. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art 3D models and standard MIL approaches. By eliminating the reliance on dense annotations, Brain-Adapter provides a highly scalable and clinically viable solution for 3D acute intracranial pathology analysis.
Enriched text-guided variational multimodal knowledge distillation network (VMD) for automated diagnosis of plaque vulnerability in 3D carotid artery MRI
用于3D颈动脉MRI斑块易损性自动诊断的增强文本引导变分多模态知识蒸馏网络(VMD)
Bo Cao, Fan Yu, Mengmeng Feng, SenHao Zhang, Xin Meng, Yue Zhang, Zhen Qian, Jie Lu
机构
*
department of Radiology and Nuclear Medicine, Xuanwu Hospital, Capital Medical University(放射医学与核医学科,宣武医院,首都医科大学)
;
Beijing Key Laboratory of Magnetic Resonance Imaging and Brain Informatics(北京磁共振成像与脑信息学重点实验室)
;
Beijing United Imaging Research Institute of Intelligent Imaging(北京智能成像联合研究院)
近年来,多模态学习因能有效利用不同模态的数据特征而备受关注。放射科医生和传统3D视觉网络直接从颈动脉3D MRI图像诊断动脉粥样硬化斑块易损性都颇具挑战。临床中,放射科医生采用整合多种成像模态和领域专业知识的多模态方法评估患者状况,为多模态诊断网络的构建奠定了基础。本文提出一种利用放射科医生领域知识的有效策略,通过变分推理与多模态知识蒸馏(Variation inference and Multimodal knowledge Distillation,VMD)实现颈动脉斑块易损性的自动诊断。该方法擅长从训练数据中有限的图像标注和放射报告中利用跨模态先验知识,从而提升未标注3D MRI图像诊断网络的准确率。我们在内部收集的数据集上开展了深入实验,验证了所提VMD策略的有效性。
英文摘要
Multimodal learning has attracted much attention in recent years due to its ability to effectively utilize data features from a variety of different modalities. Diagnosing the vulnerability of atherosclerotic plaques directly from carotid 3D MRI images is relatively challenging for both radiologists and conventional 3D vision networks. In clinical practice, radiologists assess patient conditions using a multimodal approach that incorporates various imaging modalities and domain-specific expertise, paving the way for the creation of multimodal diagnostic networks. In this paper, we have developed an effective strategy to leverage radiologists' domain knowledge to automate the diagnosis of carotid plaque vulnerability through Variation inference and Multimodal knowledge Distillation (VMD). This method excels in harnessing cross-modality prior knowledge from limited image annotations and radiology reports within training data, thereby enhancing the diagnostic network's accuracy for unannotated 3D MRI images. We conducted in-depth experiments on the dataset collected in-house and verified the effectiveness of the VMD strategy we proposed.
机构
*
Ain Shams University(艾因夏姆斯大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
University of Catania(卡塔尼亚大学)
;
Universitat de Barcelona(巴塞罗那大学)
Multiphasic contrast-enhanced CT (CECT) is widely used for abdominal lesion characterization, yet it carries inherent risks of contrast-induced nephropathy, escalates acquisition burden, and heavily contributes to radiologist workload. To address these challenges, we introduce a novel multi-center benchmark for multi-organ abdominal disease diagnosis and automated radiology report generation, which learns to synthesize contrast-enhanced findings from single-phase non-contrast CT (NCCT). To support this, we curated a large-scale dataset of paired NCCT-CECT studies and their corresponding contrast-enhanced radiology reports from two centers, partitioned into internal sets and an external validation cohort. Under a unified evaluation protocol, we benchmarked five contemporary deep learning architectures encompassing chest-specific, abdomen-specific, and general-purpose multimodal domains. Extensive experiments demonstrate that NCCT retains diagnostic signals, achieving an average multi-organ AUC of 69.1% on the internal cohort and 63.1% on the external cohort, respectively. By releasing this dataset and standardized benchmark publicly, this study aims to catalyze future research into safer, resource-efficient, and globally accessible contrast-free abdominal imaging workflows. Code is available at: https://github.com/xmed-lab/TriALS-Report.
Cyclic 2.5D Perceptual Loss for Cross-Modal 3D Medical Image Synthesis: T1w MRI to Tau PET
循环2.5D感知损失用于跨模态3D医学图像合成:T1加权MRI到tau PET
Junho Moon, Symac Kim, Haejun Chung, Ikbeom Jang
机构
*
Department of Artificial Intelligence, Hanyang University, Seoul, South Korea(人工智能系,翰阳大学,首尔,韩国)
;
Department of Electronic Engineering, Hanyang University, Seoul, South Korea(电子工程系,翰阳大学,首尔,韩国)
;
Division of Computer Engineering, Hankuk University of Foreign Studies, Yongin, South Korea(计算机工程系,韩国外语大学, Yongin,韩国)
;
Division of AI Data Convergence, Hankuk University of Foreign Studies, Yongin, South Korea(AI数据融合系,韩国外语大学, Yongin,韩国)
;
Division of Language & AI, Hankuk University of Foreign Studies, Seoul, South Korea(语言与AI系,韩国外语大学,首尔,韩国)
Positron emission tomography (PET) provides molecular biomarkers for Alzheimer's disease and related dementias (ADRD) and is increasingly used for diagnosis, staging, and clinical trial enrichment. However, its use is limited by cost, regulatory restrictions, and the invasiveness of radiotracer injection. Although current frameworks emphasize multimodal biomarker assessment, including the amyloid/tau/neurodegeneration (A/T/N) scheme, these barriers constrain access to PET imaging. Cross-modal image synthesis may help address this gap by reconstructing unavailable modalities from routine scans. Because PET is clinically valuable for regional uptake patterns rather than exact voxel-wise intensities, perceptual losses that capture higher-level semantic features are well suited to PET synthesis. Existing 2D, 3D, and 2.5D perceptual losses for 3D synthesis each have limitations, including restricted volumetric context, scarcity of pretrained 3D models, and difficulty balancing optimization across anatomical planes. In this study, we synthesize tau PET from structural MRI by generating 3D pseudo-[18F]flortaucipir standardized uptake value ratio (SUVR) maps from 3D T1-weighted MR images. We propose a cyclic 2.5D perceptual loss that alternates optimization across axial, coronal, and sagittal planes during training to improve volumetric consistency. We also standardize PET SUVRs by scanner manufacturer, reducing inter-manufacturer variability and better preserving high-uptake regions. Using cohorts spanning the ADRD spectrum from the ADNI and the SCAN cohort, we show that the method generalizes across U-Net, UNETR, SwinUNETR, CycleGAN, and Pix2Pix, with strong performance. Notably, it improves agreement between synthesized SUVRs and measured PET in brain regions relevant to Alzheimer-type tau pathology. Code is publicly available at https://github.com/labhai/Cyclic-2.5D-Perceptual-Loss.
Physically Aware Radiomics Without Interpolation: Disentangling Voxel Geometry and Signal Modification in CT and MRI
无插值的物理感知放射组学:解析CT和MRI中的体素几何与信号修正
David Corral Fontecha, Juan Miranda Bautista, Pablo Menendez Fernández-Miranda, Sergio Rubio-Martín, Lara Lloret Iglesias, Jose A. Vega
机构
*
Complejo Asistencial Universitario de León(莱昂大学附属医院综合体)
;
Hospital Universitario Rey Juan Carlos(雷·胡安·卡洛斯大学医院)
;
Health Research Institute of the Jiménez Díaz Foundation(希门尼斯·迪亚斯基金会健康研究所)
;
Rey Juan Carlos University(雷·胡安·卡洛斯大学)
;
Universidad de Oviedo(奥维耶多大学)
;
Universidad de León(莱昂大学)
;
IFCA-CSIC(西班牙国家研究委员会高级计算与电子科学研究所)
;
Universidad Autónoma de Chile(智利大学)
Objective: Radiomic texture features are usually computed in voxel-index neighborhoods, implicitly assuming isotropic spatial relationships. In anisotropic images, this can confound voxel geometry with interpolation-induced signal changes. We developed a voxel-spacing-aware radiomic framework that incorporates physical geometry into texture computation without resampling.
Approach: We modified PyRadiomics to account for voxel spacing while preserving the native image signal. Four configurations were compared: native non-resampled extraction (NR), isotropic resampling (RS), voxel-spacing-aware extraction (VS), and fake-isotropic preprocessing (FK), in which spacing metadata were overwritten without altering the image array. Experiments included 685 LIDC-IDRI pulmonary nodules and 209 I-SPY2 breast MRI cases, with 196 radiomic descriptors. Robustness was assessed using ICC, within-subject variability, Friedman testing, feature selection, machine learning, a multilayer perceptron, and external validation.
Main results: VS showed near-native agreement with NR: median ICC(A,1) was 0.9976 in CT and 0.9984 in MRI. RS produced lower agreement and larger deviations, while FK showed intermediate behavior, confirming that spacing metadata alone can affect radiomic features. Gradient-derived and neighborhood-sensitive descriptors were most affected by preprocessing. VS preserved predictive performance comparable to NR in external CT validation, whereas MRI showed greater variability across preprocessing strategies and classifiers.
Significance: Voxel-spacing-aware extraction separates geometric modeling from interpolation-induced signal modification while preserving the native image signal, offering a coherent alternative to isotropic resampling for radiomic analysis of anisotropic CT and MRI.
Multimodal CT-MRI registration is central to image-guided radiotherapy, surgical navigation, and diagnostic workflows, but most pipelines report only aggregate quality metrics without per-case reliability signals. We propose a reliability-aware framework that converts registration quality into Green/Yellow/Red risk categories using data-learned thresholds. CT images were registered to T1-weighted MRI using rigid and affine transformations on 90 paired slices from 18 patients across brain, abdominal, and neck anatomies. Reliability was assessed using Delta NMI, Delta SSIM, Dice overlap, registration stability, and inverse consistency error, combined into a single score R. Thresholds learned from training patients were applied unchanged to held-out test patients. Affine registration outperformed rigid registration on NMI and SSIM, yielding 44% Green classifications versus 33% for rigid. Reliability-filtered registrations improved the average alignment profile compared with unfiltered methods. Per-anatomy analysis showed substantial variation, with stronger reliability for abdominal registrations than brain registrations. Weight sensitivity analysis identified Dice overlap as the dominant reliability component. The proposed framework provides an interpretable quality-control layer for multimodal registration, while risk thresholds reflect statistical rather than clinical validation.
pyMEAL: A Multi-Encoder Augmentation-Aware-Learning Toolbox for Robust Medical Image Translation
pyMEAL:用于稳健医学图像翻译的多编码器增强感知学习工具箱
Abdul-mojeed Olabisi Ilyas, Adeleke Maradesa, Jamal Banzi, Jianpan Huang, Henry K. F. Mak, Kannie W. Y. Chan
机构
*
Hong Kong Centre for Cerebro-Cardiovascular Health Engineering (COCHE)(香港脑心血管健康工程中心)
;
Department of Informatics, Sokoine University of Agriculture(农业科学学院信息系)
;
Department of Diagnostic Radiology, The University of Hong Kong(香港大学诊断放射科)
;
State Key Laboratory of Brain and Cognitive Sciences, The University of Hong Kong(香港大学脑科学与认知科学国家重点实验室)
;
Alzheimer’s Disease Research Network, The University of Hong Kong(香港大学阿尔茨海默病研究网络)
;
Department of Biomedical Engineering, City University of Hong Kong(城市大学生物医学工程系)
;
Department of Radiology and Radiological Science, The Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院放射学与放射科学系)
;
City University of Hong Kong Shenzhen Research Institute(城市大学深圳研究 institute)
Medical imaging plays a vital role in clinical diagnosis, yet AI-driven imaging methods remain challenged by patient variability, image artifacts, and limited robustness across acquisition conditions. Although deep learning has advanced medical image analysis, 3D image translation remains hindered by limited training data and variability arising from scanner differences, imaging protocols, and patient motion. Conventional data augmentation typically relies on a single transformation pipeline, overlooking augmentation-specific characteristics and limiting representation learning.
To address these challenges, we propose Multi-Encoder Augmentation-Aware Learning (MEAL), which processes multiple augmentation variants through dedicated encoder pathways. Three feature integration strategies are investigated: encoder concatenation (MEAL-CC), fusion layer (MEAL-FL), and an adaptive controller block (MEAL-BD). By dynamically weighting augmentation-specific features before decoding, MEAL-BD preserves complementary representations and improves robustness to clinically relevant variability.
We evaluate MEAL using CT-to-T1-weighted MRI translation, a clinically relevant task when MRI is unavailable, contraindicated, or delayed. Across predefined and unseen test datasets, MEAL-BD consistently outperformed competing approaches under both geometric perturbations and standard imaging conditions, achieving higher peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM). By prioritizing structural fidelity over perceptual realism, MEAL supports clinical interpretation and downstream image analysis rather than replacing diagnostic MRI, demonstrating that augmentation-aware representation learning improves the robustness and clinical applicability of medical image translation.
WING: A Window-Prior-Based Generative Network with Gated Inception for Cross-Modality CT Synthesis
WING:一种基于窗口先验的带门控inception的生成网络用于跨模态CT合成
Siyuan Mei, Yan Xia, Yipeng Sun, Siming Bayer, Zirong Li, Chengze Ye, Daiqi Liu, Fuxin Fan, Yixing Huang, Andreas Maier
机构
*
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg(模式识别实验室,埃尔朗根-纽伦堡弗里德里希-亚历山大大学)
;
Department of Orthodontics and Orofacial Orthopaedics, Friedrich-Alexander-Universität Erlangen-Nürnberg(正畸与口腔颌面正畸科,埃尔朗根-纽伦堡弗里德里希-亚历山大大学)
;
Siemens Healthineers(西门子医疗)
;
Institute of Medical Technology, Peking University(北京大学医学技术研究所)
Generating CT volumes from MRI and CBCT can improve treatment planning in adaptive radiotherapy while avoiding additional radiation exposure. However, direct regression of CT intensities is challenged by the inherently high dynamic range and long-tailed distributions, thereby averaging out sparse yet clinically important structures. To alleviate this issue, we reformulate the regression target into multiple windowed representations, leveraging the inductive prior that CT intensities are structure-deterministic and window-separable. These windowed views exhibit smoother distributions and admit structured fusion back to the full-range CT. Building on this reformulation, we introduce WING, a WINdow-prior-based Generative network comprising: 1) a new Gated Inception Generator to produce multi-window predictions, enabling multi-shape kernel interactions to capture cross-modality correspondence; 2) a Fuse-and-Refine Transformer to aggregate the windowed outputs and learn residuals for detail refinement; and 3) a joint adversarial training objective to enhance window-conditioned realism. Extensive experiments demonstrate that our compact WING achieves state-of-the-art performance on the MRI-to-CT and CBCT-to-CT benchmarks, while supporting multi-anatomy synthesis with a single model.
This paper presents and validates CTseg, a freely available software for brain CT segmentation, spatial normalisation, and volumetrics. CTseg builds on the Multi-Brain generative modelling framework, providing a CT-specific pipeline that produces tissue maps, deformation fields, and brain volume estimates in the same format as SPM's unified segmentation, thereby extending SPM's established analysis chain from MRI to CT. CTseg is designed for routine hospital CT scans without requiring preprocessing or resampling in deployment. Although CTseg has been adopted in clinical research spanning, among other things, stroke, dementia, and brain morphometry, a systematic validation against an independent reference standard has been lacking. Using paired MR/CT head scans, we evaluate CTseg across four dimensions: segmentation accuracy against an MRI-derived silver standard; spatial normalisation consistency through group-average sharpness and voxelwise coefficient of variation; brain volume agreement via intraclass correlation and Bland-Altman analysis; and downstream sex classification performance from normalised tissue maps. As a baseline, we apply SPM's MRI-based unified segmentation directly to the CT images. CTseg significantly outperformed this baseline for segmentation and normalisation, showed stronger TBV agreement, and achieved comparable TIV agreement. CTseg is freely available at https://github.com/WCHN/CTseg, and all experiment code is included in the repository for full reproducibility.
CommentsPreprint manuscript, 16 pages, 4 figures, 3 tables. This manuscript presents a residual-bounded 2.5D CT/CTA restoration framework for conservative medical image enhancement and evaluates it using image-recovery, baseline comparison, Monte Carlo stability, anatomical localization, and external low-dose CT testing
Image restoration models are increasingly applied to degraded medical scans, but in safety-sensitive settings they must improve image quality without uncontrolled modification of clinically important regions. This is especially relevant for intracranial CT and CT angiography (CTA), where small vessels and aneurysm-relevant cues lie near high-contrast anatomical boundaries. We frame medical image restoration as a conservative AI problem and present a residual-bounded 2.5D restoration framework trained on synthetically degraded CT/CTA inputs. The model adds a learned residual to the original center slice through an edit-control map that limits the magnitude and spatial extent of modification. We evaluate the framework using an aneurysm-relevant image-recovery matrix, paired comparison against a Gaussian baseline, Monte Carlo stability testing, anatomical localization of meaningful edits, and external evaluation on low-dose CT. On 50 out-of-distribution CT-CTA cases, the bounded model achieved a mean target gain of 0.0635, a mean PSNR of 37.51 dB, and an iatrogenic-edit rate of 4.0%. Across 1,000 Monte Carlo runs, it remained net positive in 85.4% of runs with no stably negative cases. On external low-dose CT, the model was directionally beneficial and produced a substantially smaller modification footprint than the baseline. Meaningful edits concentrated in brain and skull regions while unrelated anatomy showed negligible change. These findings provide preliminary computational evidence that residual-bounded restoration is feasible in boundary-sensitive vascular imaging, but they do not establish clinical diagnostic performance and require expert review and prospective validation before clinical use.
AI Approach for MRI-only Full-Spine Vertebral Segmentation and 3D Reconstruction in Paediatric Scoliosis
基于AI的MRI-only全脊柱椎体分割及3D重建在儿童脊柱侧弯中的应用
Nathasha Naranpanawa, Maree T. Izatt, Robert D. Labrom, Geoffrey N. Askin, J. Paige Little
机构
*
Biomechanics and Spine Research Group, School of Mechanical, Medical, and Process Engineering, Faculty of Engineering, Queensland University of Technology(生物力学与脊柱研究组,机械、医学与过程工程学院,工程学院,昆士兰理工大学)
;
Centre for Biomedical Technologies, School of Mechanical, Medical and Process Engineering, Faculty of Engineering, Queensland University of Technology(生物医学技术中心,机械、医学与过程工程学院,工程学院,昆士兰理工大学)
;
Orthopaedics Department, Queensland Children’s Hospital(骨科部门,昆士兰儿童医院)
MRI is preferred over CT in paediatric imaging because it avoids ionising radiation, but its use in spine deformity assessment is largely limited by the lack of automated, high-resolution 3D bony reconstruction, which continues to rely on CT. MRI-based 3D reconstruction remains impractical due to manual workflows and the scarcity of labelled full-spine datasets. This study introduces an AI framework that enables fully automated thoracolumbar spine (T1-L5) segmentation and 3D reconstruction from MRI alone. Historical low-dose CT scans from adolescent idiopathic scoliosis (AIS) patients were converted into MRI-like images using a GAN and combined with existing labelled thoracic MRI data to train a U-Net-based model. The resulting algorithm accurately generated continuous thoracolumbar 3D reconstructions, improved segmentation accuracy (88% Dice score), and reduced processing time from approximately 1 hour to under one minute, while preserving AIS-specific deformity features. This approach enables radiation-free 3D deformity assessment from MRI, supporting clinical evaluation, surgical planning, and navigation in paediatric spine care.
The deployment of artificial intelligence in medical imaging is hindered by high computational complexity and resource-intensive processing of volumetric data. Although chest computed tomography (CT) volumes offer richer diagnostic information than projection radiography, their use in AI-based diagnosis remains limited due to the computational burden of processing uncompressed volumetric images (typically stored in NIfTI or DICOM format). Addressing the growing need for low-resource deployment and efficient electronic data transfer, we investigate the utilization of JPEG-compressed chest CT volumes for thoracic abnormality detection. We propose Feature Attention Style Transfer (FAST), a novel distillation framework that transfers both activation patterns and structural relationships from high-fidelity CT representations to a spatiotemporal visual encoder operating on compressed inputs. By combining Gram-matrix-based attention style preservation with dual-attention feature alignment, FAST enables robust feature extraction from degraded volumes. Furthermore, we introduce Structured Factorized Projection (SFP), leveraging Block Tensor Train decomposition as a parameter-efficient alternative to dense projection layers, reducing projection-head parameters by almost half. Our contrastive learning pipeline, CT-Lite, integrates these components with a SigLIP-based multimodal alignment objective. Experiments on CT-RATE, NIDCH, and Rad-ChestCT demonstrate that CT-Lite achieves AUROC within 5-7\% of the uncompressed-input baseline across all three datasets, despite operating on compressed inputs with significantly fewer parameters, paving the way for AI-based clinical evaluation under resource constraints.
NeuroBridge: Bridging Multi-Task MRI Knowledge for Neurodegenerative Disease Diagnosis
NeuroBridge:桥接多任务MRI知识用于神经退行性疾病诊断
Mengyu Li, Guoyao Shen, Chad W. Farris, Xin Zhang
机构
*
Department of Mechanical Engineering, Boston University(波士顿大学机械工程系)
;
The Photonics Center, Boston University(波士顿大学光子中心)
;
Rafik B. Hariri Institute for Computing and Computational Science & Engineering, Boston University(波士顿大学拉菲克·B·哈里里计算与计算科学与工程研究所)
;
Department of Radiology, Boston University Chobanian & Avedisian School of Medicine(波士顿大学切博尼亚与阿维迪亚安医学学院放射科)
;
Department of Radiology, Boston Medical Center(波士顿医学中心放射科)
;
Department of Electrical and Computer Engineering, Boston University(波士顿大学电气与计算机工程系)
;
Department of Biomedical Engineering, Boston University(波士顿大学生物医学工程系)
;
Division of Materials Science and Engineering, Boston University(波士顿大学材料科学与工程系)
INTRODUCTION: Accurate MRI-based identification of Alzheimer's disease (AD), mild cognitive impairment (MCI), and related dementias remains challenging because disease-related structural changes are often subtle and heterogeneous. We developed NeuroBridge, a clinically guided multi-task MRI framework for neurodegenerative disease diagnosis. METHODS: NeuroBridge integrates large-scale self-supervised MRI pretraining with hippocampal segmentation, hippocampal atrophy classification, and reconstruction objectives, followed by gated fusion fine-tuning. Performance was evaluated across ADNI and OASIS cohorts, including cross-cohort transfer, probability-based analysis, and opportunistic screening. RESULTS: NeuroBridge achieved the highest performance across evaluated classification tasks, reaching 88.17% accuracy for AD versus cognitively normal controls in ADNI and 82.78% in OASIS. The largest gains occurred in MCI-related and mixed-diagnosis settings. The framework demonstrated strong cross-cohort generalization, systematic associations between predicted-class probability and accuracy, and the feasibility of probability-based opportunistic screening. DISCUSSION: Clinically guided multi-task representation learning improves neurodegenerative MRI diagnosis beyond conventional single-task approaches. NeuroBridge provides a robust and scalable framework for dementia assessment and MRI-based opportunistic screening.
Self-Auditing Residual Drifting for Pathology-Preserving Accelerated Knee MRI
自审计残差漂移模型用于保留病理特征的加速膝关节MRI
Qing Lyu, Jianxu Wang, Mohammad Kawas, Ge Wang, Christopher T. Whitlow
机构
*
Department of Radiology & Biomedical Imaging, Yale School of Medicine(放射学与生物医学成像系,耶鲁医学院)
;
Department of Biomedical Engineering, Rensselaer Polytechnic Institute(生物医学工程系,伦塞拉尔理工学院)
Accelerated magnetic resonance imaging reduces acquisition time, but reconstruction from undersampled k-space can blur diagnostically relevant structures or introduce failures that are not captured by global image metrics. We propose SA-RDM-DC, a Self-Auditing Residual generative Drifting Model with Data Consistency for accelerated knee MRI. The method adapts the newly proposed generative drifting paradigm to accelerated MRI by training a physics-conditioned drift field from the zero-filled reconstruction toward the fully sampled residual correction. It predicts image- and missing-k-space residual corrections, enforces data consistency with acquired k-space, uses frequency-aware and residual drifting supervision to recover fine detail, and produces dense error maps and slice-level risk scores in the same inference pass. We evaluate SA-RDM-DC on multi-coil fastMRI knee data at acceleration factors of 4, 8, and 12, with fastMRI+ pathology annotations for region-level and classifier-based task preservation, and on SKM-TEA for zero-shot and fine-tuned protocol-shift evaluation. Compared with zero-filled reconstruction, UNet-image-SENSE, DC-UNet, Score-Diffusion, ELF-Diff, SENSE-VarNet, and MoDL baselines, SA-RDM-DC achieves the highest SSIM across fastMRI acceleration factors while retaining subsecond per-slice inference and avoiding the long sampling time of iterative diffusion baselines. In pathology-aware analysis, SA-RDM-DC preserves lesion-region structural fidelity and reduces meniscus prediction instability. Its self-auditing scores strongly identify high-error reconstructions on fastMRI and partially transfer as a selective-review signal under SKM-TEA protocol shift. These results support reconstruction evaluation that jointly considers image fidelity, pathology preservation, runtime, and case-specific reliability.
7 Tesla Quantitative MRI and Machine Learning for Exploratory Motor Subtype Stratification and Diagnosis in Parkinson's Disease
7特斯拉定量MRI与机器学习用于帕金森病运动亚型分层和探索性诊断
Anne Louise Kristoffersen, Runa Geirmundsdatter Unsgård, Marc-Antoine Fortin, Ingrid Gylterud Kvålsgard, Kjersti Eline Stige, Thanh Pierre Doan, Erik Magnus Berntsen, Charalampos Tzoulis, Pål Erik Goa
AI总结
本研究利用7特斯拉MRI的定量图和深度学习自动脑分割,结合特征选择的机器学习分类器,实现了帕金森病运动亚型(姿势不稳与步态困难型 vs 震颤为主型)的高精度分层和诊断。
详情
AI中文摘要
帕金森病(PD)是一种高度异质性疾病,包括哪些运动症状占主导。支持亚型分层的影像生物标志物可以改善生物学理解和研究设计,并实现个性化治疗策略。本研究评估了基于深度学习的自动脑分割,结合7特斯拉MRI的定量图,是否能突出健康对照(HC)、姿势不稳与步态困难(PIGD)和震颤为主(TD)之间的差异,并随后用于客观的PD分层。特征选择可能提高机器学习分类器的性能。研究纳入21名HC和24名PD患者(PwP)。U-Net训练通过DSC评估。定义了两种分类方法,使用5折交叉验证,针对三个任务:(1)HC vs PwP;(2)PIGD vs TD;(3)多类,HC vs PIGD vs TD。方法A使用所有提取的特征。方法B找到分类任务的最优特征子集。U-Net在训练期间对所有ROI的平均DSC为0.86。方法A:任务1最佳准确率0.69,最佳AUC 0.73;任务2准确率0.69,AUC 0.90;任务3准确率0.62,AUC 0.66。方法B:任务1准确率0.82,AUC 0.93;任务2准确率1.00,AUC 1.00;任务3准确率0.73,AUC 0.91。基于深度学习的分割结合qMRI特征选择相对于使用所有特征提高了分类性能,支持可解释的低维成像特征用于PD诊断支持和表型分层的潜力。需要更大规模的多中心研究来评估泛化性和稳定性。
英文摘要
Parkinson's disease (PD) is a highly heterogeneous disease, including which motor symptoms are dominating. Imaging biomarkers that support subtype stratification could also improve biological understanding and study design, and enable personalized treatment strategies. This study evaluates whether deep-learning based automatic brain segmentation, in addition to quantitative maps from 7 Tesla MRI, can highlight differences between Healthy Controls (HC), Postural Instability and Gait Difficulty (PIGD) and Tremor Dominant (TD), and subsequently be used for objective PD stratification. The performance of machine learning classifiers may be improved with feature selection. 21 HC, and 24 people with PD (PwP) were included. The U-Net training was assessed with DSC. Two classification approaches using 5-fold cross-validation were defined across three tasks: (1) HC vs PwP; (2) PIGD vs TD; (3) multiclass, HC vs PIGD vs TD. Approach A used all extracted features. Approach B found the optimal subset of features for the classification tasks. The U-Net achieved mean DSC of 0.86 for all ROIs during training. Approach A: Task 1 best accuracy of 0.69 and best AUC of 0.73. Task 2 accuracy 0.69, AUC 0.90. Task 3 accuracy 0.62, AUC 0.66. Approach B: Task 1 accuracy of 0.82 and AUC of 0.93. Task 2 accuracy 1.00, AUC 1.00. Task 3 accuracy 0.73, AUC 0.91. DL-based segmentation combined with qMRI feature selection improved classification relative to using all features, supporting the potential of interpretable, low-dimensional imaging signatures for PD diagnosis support and phenotype stratification. Larger, multi-site studies are warranted to assess generalizability and stability.
While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge when training on 3D-CT volumetric data. In this study, we propose TotalFM, a radiological foundation model that efficiently learns the correspondence between 3D-CT images and linguistic expressions based on the concept of organ separation, utilizing a large-scale dataset of 140,000 series. By automating the creation of organ volume and finding-sentence pairs through segmentation techniques and Large Language Model (LLM)-based radiology report processing, and by combining self-supervised pre-training via VideoMAE with contrastive learning using volume-text pairs, we aimed to balance computational efficiency and representation capability. In zero-shot organ-wise lesion classification tasks, the proposed model achieved higher F1 scores in 83% (5/6) of organs compared to CT-CLIP and 64% (9/14) of organs compared to Merlin. These results suggest that the proposed model exhibits high generalization performance in a clinical evaluation setting using actual radiology report sentences. Furthermore, in zero-shot finding-wise lesion classification tasks, our model achieved a higher AUROC in 83% (25/30) of finding categories compared to Merlin. We also confirmed performance comparable to existing Vision-Language Models (VLMs) in radiology report generation tasks. Our results demonstrate that the organ-separated learning framework can serve as a realistic and effective design guideline for the practical implementation of 3D-CT foundation models. The source code and pretrained models are publicly available at https://github.com/jichi-labo/TotalFM.
Automated 3D radiology report generation often suffers from clinical hallucinations and a lack of the iterative verification found in human practice. While recent Vision-Language Models (VLMs) have advanced the field, they typically operate as monolithic "black-box" systems without the collaborative oversight characteristic of clinical workflows. To address these challenges, we propose MARCH (Multi-Agent Radiology Clinical Hierarchy), a multi-agent framework that emulates the professional hierarchy of radiology departments and assigns specialized roles to distinct agents. MARCH utilizes a Resident Agent for initial drafting with multi-scale CT feature extraction, multiple Fellow Agents for retrieval-augmented revision, and an Attending Agent that orchestrates an iterative, stance-based consensus discourse to resolve diagnostic discrepancies. On the RadGenome-ChestCT dataset, MARCH significantly outperforms state-of-the-art baselines in both clinical fidelity and linguistic accuracy. Our work demonstrates that modeling human-like organizational structures enhances the reliability of AI in high-stakes medical domains.
机构
*
School of Biomedical Engineering, Tsinghua University(清华大学生物医学工程学院)
;
Department of Radiology, Mianyang Central Hospital(绵阳市中心医院放射科)
;
Center for Biomedical Imaging Research, Tsinghua University(清华大学生物医学成像研究中心)
Chest computed tomography (CT) is central to the detection and management of thoracic disease, yet the growing scale and complexity of volumetric imaging increasingly exceed what can be addressed by scan-level prediction alone. Clinically useful AI for CT must not only recognize disease across the whole volume, but also localize abnormalities and provide interpretable visual evidence. Existing vision-language foundation models typically compress scans and reports into global image-text representations, limiting their ability to preserve spatial evidence and support clinically meaningful interpretation. Here we developed EXACT, an explainable anomaly-aware foundation model for three-dimensional chest CT that learns spatially resolved representations from paired clinical scans and radiology reports. EXACT was pre-trained on 25,692 CT-reports pairs using anatomy-aware weak supervision, jointly learning organ segmentation and multi-instance anomaly localization without manual voxel-level annotations. The resulting organ-specific anomaly-aware maps assign each voxel a disease-specific anomaly score confined to its corresponding anatomy, jointly encoding lesion extent and organ-level context. In retrospective multinational and multi-center evaluations, EXACT showed broad and consistent improvements across clinically relevant CT tasks, spanning multi-disease diagnosis, zero-shot anomaly localization, downstream adaptation, and visually grounded report generation, outperforming existing three-dimensional medical foundation models. By transforming routine clinical CT scans and free-text reports into explainable voxel-level representations, EXACT establishes a scalable paradigm for trustworthy volumetric medical AI.
Generative Drifting for Conditional Medical Image Generation
生成漂移用于条件医学图像生成
Zirong Li, Siyuan Mei, Weiwen Wu, Andreas Maier, Lina Gölz, Yan Xia
机构
*
Department of Orthodontics and Orofacial Orthopedics, Friedrich-Alexander-University Erlangen-Nuremberg(口腔正畸与面部骨科系,弗里德里希-艾萨克-埃尔兰-纽伦堡大学)
;
Department Artificial Intelligence in Biomedical Engineering, Friedrich-Alexander-University Erlangen-Nuremberg(生物医学工程人工智能系,弗里德里希-艾萨克-埃尔兰-纽伦堡大学)
;
Pattern Recognition Lab, Friedrich-Alexander-University Erlangen-Nuremberg(模式识别实验室,弗里德里希-艾萨克-埃尔兰-纽伦堡大学)
;
Department of Biomedical Engineering, Sun-Yat-sen University(生物医学工程系,孙中山大学)
Conditional medical image generation plays an important role in many clinically relevant imaging tasks. However, existing methods still face a fundamental challenge in balancing inference efficiency, patient-specific fidelity, and distribution-level plausibility, particularly in high-dimensional 3D medical imaging. In this work, we propose GDM, a generative drifting framework that reformulates deterministic medical image prediction as a multi-objective learning problem to jointly promote distribution-level plausibility and patient-specific fidelity while retaining one-step inference. GDM extends drifting to 3D medical imaging through an attractive-repulsive drift that minimizes the discrepancy between the generator pushforward and the target distribution. To enable stable drifting-based learning in 3D volumetric data, GDM constructs a multi-level feature bank from a medical foundation encoder to support reliable affinity estimation and drifting field computation across complementary global, local, and spatial representations. In addition, a gradient coordination strategy in the shared output space improves optimization balance under competing distribution-level and fidelity-oriented objectives. We evaluate the proposed framework on two representative tasks, MRI-to-CT synthesis and sparse-view CT reconstruction. Experimental results show that GDM consistently outperforms a wide range of baselines, including GAN-based, flow-matching-based, and SDE-based generative models, as well as supervised regression methods, while improving the balance among anatomical fidelity, quantitative reliability, perceptual realism, and inference efficiency. These findings suggest that GDM provides a practical and effective framework for conditional 3D medical image generation.
Plug-and-Play diffusion prior (PnPDP) frameworks have emerged as a powerful paradigm for solving imaging inverse problems by treating pretrained generative models as modular priors. However, we identify a critical flaw in prevailing PnP solvers (e.g., based on HQS or Proximal Gradient): they function as memoryless operators, updating estimates solely based on instantaneous gradients. This lack of historical tracking inevitably leads to non-vanishing steady-state bias, where the reconstruction fails to strictly satisfy physical measurements under heavy corruption. To resolve this, we propose Dual-Coupled PnP Diffusion (DC-PnPDP), which restores the classical dual variable to provide integral feedback, progressively enforce agreement between the data-consistency and prior. However, this rigorous geometric coupling introduces a secondary challenge: the accumulated dual residuals exhibit spectrally colored, structured artifacts that violate the Additive White Gaussian Noise (AWGN) assumption of diffusion priors, causing severe hallucinations. To bridge this gap, we introduce Spectral Homogenization (SH), a frequency-domain adaptation mechanism that modulates these structured residuals into statistically compliant pseudo-AWGN inputs. This effectively aligns the solver's rigorous optimization trajectory with the denoiser's valid statistical manifold. Extensive experiments on CT and MRI reconstruction demonstrate that our approach resolves the bias-hallucination trade-off, achieving state-of-the-art fidelity with significantly accelerated convergence. The code is available at https://github.com/duchenhe/DC-PnPDP
Prostate cancer (PCa) is one of the most common cancers in men worldwide. Bi-parametric MRI (bp-MRI) and clinical variables are crucial for PCa identification and improving treatment decisions. However, this process is subjective to expert interpretations. Furthermore, most existing computer-aided diagnosis methods focus on imaging-based models, overlooking the clinical context and suffering from data scarcity, limiting their ability to learn robust representations. We propose a geometric multimodal Foundation Model (FM), named MFM-Geom, that learns representations from bp-MRI and clinical reports, encoding visual findings and information from the context of clinical variables. In the representations classification head, the approach leverages symmetric positive definite (SPD) matrices and Riemannian deep learning to integrate imaging-text representations from a biomedical multimodal FM. Using 10% of the training data, MFM-Geom outperformed baseline class token embedding-based classification (+8.3%, AUC-PR of 90.67). Generalization on external dataset confirmed the robustness of fine-tuning biomedical FM, achieving an AUC-PR of 90.6.
CORA: Generalizable coronary artery disease assessment and risk stratification from coronary CT angiography using pathology-centric representation learning
CORA:一种由病理合成驱动的冠状动脉CT血管造影分析和MACE风险评估基础模型
Jinkui Hao, Gorkem Durak, Halil Ertugrul Aktas, Ulas Bagci, Bradley D. Allen, Nilay S. Shah, Bo Zhou
机构
*
Department of Radiology, Northwestern University(西北大学放射学系)
;
Department of Biomedical Engineering, Northwestern University(西北大学生物医学工程系)
;
Department of Cardiology, Northwestern University(西北大学心脏病学系)
;
Department of Preventive Medicine, Northwestern University(西北大学预防医学系)
Coronary artery disease, a leading cause of cardiovascular mortality worldwide, can be assessed non-invasively by coronary computed tomography angiography (CCTA). Although deep learning has advanced automated CCTA analysis, clinical translation remains constrained by the scarcity of expert-annotated data and by the spatial sparsity of coronary pathology, which occupies only a small fraction of each scan. Widely used label-free pretraining strategies, such as masked image modeling and contrastive learning, optimize for global anatomical reconstruction and tend to under-represent these tiny localized pathological features. Here we present CORA, an annotation-efficient model for comprehensive coronary artery disease assessment. Rather than reconstructing background anatomy, CORA learns from volumetric CCTA through a synthesis-driven self-supervised strategy: an anatomy-guided engine inserts diverse synthetic calcified and non-calcified lesions into unlabeled scans, reframing pretraining as an abnormality-detection task that biases representation learning toward clinically relevant disease features. We pretrained CORA on 10,138 unlabeled CCTA volumes and evaluated it across datasets from nine independent hospitals. Across plaque characterization, stenosis detection, and coronary artery segmentation, CORA consistently outperformed strong self-supervised pretraining baselines, with the largest gains on external multi-center data, indicating robust generalization under distributional shift. Coupling the imaging encoder with structured clinical variables further enabled near-term major adverse cardiac event (MACE) risk stratification. Our results show that pathology-centric, synthesis-driven pretraining is an effective and scalable strategy for annotation-efficient coronary artery disease assessment from CCTA.
HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT
头颈肿瘤 (HECKTOR) 2025 挑战赛:多模态 PET/CT 中的分割、诊断与预后基准
Numan Saeed, Salma Hassan, Shahad Hardan, Lishan Cai, Xinglong Liang, Moona Mazher, Abdul Qayyum, Yansong Bu, Mengye Lyu, Yue Lin, Mingyuan Meng, Chuanyi Huang, Lisheng Wang, Dalal Chamseddine, Shamimeh Ahrari, Beining Wu, Yifei Chen, Fuyou Mao, Hao Zhang, Baixiang Zhao, Surajit Ray, Muzi Guo, Lei Xiang, Jakob Dexl, Michael Ingrisch, Adrien Depeursinge, Arman Rahmim, Mathieu Hatt, Vincent Andrearczyk, Mohammad Yaqub
机构
*
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
Amsterdam UMC(阿姆斯特丹大学医学中心)
;
The Netherlands Cancer Institute(荷兰癌症研究所)
;
Radboud University Medical Centre(拉德堡德大学医学中心)
;
University College London(伦敦大学学院)
;
Imperial College London(帝国理工学院)
;
Shenzhen Technology University(深圳技术大学)
;
Shenzhen University(深圳大学)
;
Newland Digital Technology(新大陆数字技术)
;
The University of Sydney(悉尼大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University Hospital, Nantes(南特大学医院)
;
Nantes Université, Centrale Nantes, CNRS, LS2N(南特大学、南特中央理工学院、法国国家科学研究中心、LS2N实验室)
;
Hangzhou Dianzi University(杭州电子科技大学)
;
Tsinghua University(清华大学)
;
Central South University(中南大学)
;
University of Glasgow(格拉斯哥大学)
;
China Mobile System Integration Co., Ltd.(中移系统集成有限公司)
;
Subtle Medical Inc.(Subtle Medical公司)
;
University Hospital, LMU Munich(慕尼黑大学医院)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
BC Cancer Research Institute(不列颠哥伦比亚癌症研究所)
;
HES-SO Valais-Wallis University of Applied Sciences and Arts(HES-SO瓦莱州应用科学与艺术大学)
;
Lausanne University Hospital (CHUV)(洛桑大学医院)
;
LaTIM, INSERM, UMR 1101, Univ Brest(LaTIM实验室、法国国家健康与医学研究院、UMR 1101、布雷斯特大学)
Comments17 pages, 4 figures, 4 tables. Overview paper for the HECKTOR 2025 challenge, held as a satellite event at MICCAI 2025. Challenge website: https://hecktor.grand-challenge.org/
Head and neck cancers (HNC) represent a significant global health burden, with accurate tumor delineation being essential for effective radiotherapy planning. The complexity of the oropharyngeal anatomy, combined with the heterogeneous appearance of tumors on imaging, makes manual segmentation time-intensive and subject to inter-observer variability. Beyond segmentation, predicting long-term clinical outcomes, such as recurrence-free survival (RFS), and determining human papillomavirus (HPV) status from noninvasive imaging, remain challenging yet clinically valuable goals. The HECKTOR 2025 challenge addresses these needs by establishing a comprehensive benchmark for automated HNC analysis using multimodal PET/CT imaging and electronic health records. Building on previous editions (2020-2022), this challenge features an expanded multi-institutional dataset comprising over 1,100 patients from 10 centers worldwide. Participants were tasked with three complementary objectives: (1) segmenting primary gross tumor volumes (GTVp) and metastatic lymph nodes (GTVn), (2) predicting recurrence-free survival, and (3) classifying HPV status. The challenge attracted 35 registered teams, with 15 final submissions evaluated on a held-out test set. Top-performing algorithms achieved a mean Dice similarity coefficient of 0.75 for segmentation, a concordance index of 0.66 for survival prediction, and a balanced accuracy of 0.56 for HPV classification. This paper presents a comprehensive analysis of the submitted methodologies, evaluates their performance across different lesion characteristics, and discusses their implications for clinical translation in automated oncology workflows and decision support systems.