arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Rad-JEPA 3D:用于三维计算机断层扫描的放射学联合嵌入预测模型

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

Quoc-Huy Trinh, Minh-Van Nguyen, Ulas Bagci

arXiv 2607.26196首次发表:更新:

发表机构

Northwestern University; Technical University of Denmark(西北大学; 丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Rad-JEPA 3D是用于3D CT的自监督联合嵌入预测框架,采用混合H-Mamba编码器与HSOR正则化,在12万次CT扫描预训练后,在器官识别、空间推理等任务上达到最优性能。

AI 中文摘要

自监督预训练是三维医学图像分析的核心,未标记的CT体积数据十分丰富,但专家标注却十分稀缺。然而现有的体积编码器通常无法保留下游推理所依赖的粗空间和几何结构,当与语言模型结合时,会限制其在器官解耦、异常检测和空间理解方面的性能。我们提出Rad-JEPA 3D,这是一种联合嵌入预测框架,通过从掩码视图预测完整扫描的潜在特征来学习体积CT表示。其核心是混合H-Mamba编码器,它融合了Mamba状态空间分支(通过顺序扫描建模片间连续性)和分组查询注意力分支(捕获跨平面空间上下文),并通过轻量级的逐令牌路由器进行组合。为了提高中间表示的质量,我们进一步提出隐藏状态正交正则化(HSOR),该方法对齐学生-教师隐藏状态并减少整个编码器中的特征冗余。这种分层正则化产生更一致且具有判别力的体积表示,从而提升器官识别和空间推理任务的性能。Rad-JEPA 3D在约120,000次CT扫描上进行预训练,尽管规模紧凑(总参数仅40亿),仍取得了最先进的结果:在封闭式VQA上达到与最先进方法相当的结果,并在Spatial-Med基准上获得最佳平均空间推理得分。消融研究证实,混合模块和HSOR带来互补增益,且诱导的空间结构可在体积推理任务中替代原始语言模型的规模。

英文摘要

Self-supervised pretraining is central to 3D medical image analysis, where unlabeled CT volumes are abundant but expert annotations are scarce. Yet existing volumetric encoders often fail to preserve the coarse spatial and geometric structure that downstream reasoning depends on, limiting their performance on organ disentanglement, abnormality detection, and spatial understanding when paired with language models. We introduce Rad-JEPA 3D, a joint-embedding predictive framework that learns volumetric CT representations by predicting the latent features of a complete scan from a masked view. At its core is a hybrid H-Mamba encoder that fuses a Mamba state-space branch, which models inter-slice continuity through sequential scanning, with a grouped-query attention branch, which captures cross-plane spatial context, combined through a lightweight per-token router. To improve the quality of intermediate representations, we further propose Hidden States Orthogonal Regularization (HSOR), which aligns student-teacher hidden states and reduces feature redundancy throughout the encoder. This layer-wise regularization produces more consistent and discriminative volumetric representations, leading to improved performance on organ recognition and spatial reasoning tasks. Pretrained on approximately 120,000 CT scans, Rad-JEPA 3D attains state-of-the-art results despite its compact size: with only 4.0B total parameters, it achieves competitive results with state-of-the-art on closed-ended VQA and the best average spatial-reasoning score on the Spatial-Med benchmark. Ablation studies confirm that the hybrid block and HSOR contribute complementary gains, and that the induced spatial structure can substitute for raw language-model scale on volumetric reasoning tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑