arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HarMoE:基于数据集解耦专家的多源胸部X射线预训练

HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts

Haozhe Luo, Ziyu Zhou, Shelley Zixin Shu, Mauricio Reyes

arXiv 2608.02252首次发表:更新:

发表机构

University of Bern; Shanghai Jiao Tong University(伯尔尼大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出HarMoE框架,通过解耦数据集专家从异构胸部X射线数据集学习,提升了放射学视觉-语言模型的零样本分类、分布外迁移等性能,相关代码与87.3万份数据集将公开。

AI 中文摘要

近期用于胸部X射线理解的视觉-语言模型大多基于图像-报告对齐构建,因此严重依赖MIMIC-CXR作为主要预训练源。尽管该范式在规模上有效,但未充分探索另一重要监督源:一系列现有多标签分类数据集,这些数据集比自由文本报告提供更清晰、更明确的疾病信号,跨源组合时可提供更广泛的病理覆盖。然而,从这类异构数据集学习并非易事,因为标签本体、标注协议、采集流程和报告风格的差异可能导致模型将临床语义与数据集身份纠缠,尽管规模扩大仍会导致迁移性能不佳。本研究从协调多源学习的视角重新审视放射学视觉-语言模型的构建,提出HarMoE,这是一种感知数据集的混合专家框架,可学习跨数据集的共享医学语义,同时将源特定变化限制在更深解码器层的轻量级残差专家中。为进一步利用带标注数据集的清晰监督,我们采用统一疾病词汇与掩码多数据集监督进行训练,使模型能利用互补标注而不引入假阴性。在大规模胸部X射线基准上的实验表明,HarMoE在零样本分类、分布外迁移和定位任务上均优于强基线。我们的结果表明,构建稳健的放射学视觉-语言模型需要超越单源图像-报告对齐,转向从具有更清晰监督和更广泛覆盖的异构数据集构建结构化知识。代码及87.3万份协调数据集将在该httpsURL发布。

英文摘要

Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the dominant pretraining source. While effective at scale, this paradigm underexplores an important alternative source of supervision: a range of existing multi-label classification datasets, which provide cleaner and more explicit disease signals than free-text reports, and can offer broader pathology coverage when combined across sources. However, learning from such heterogeneous datasets is nontrivial, as differences in label ontologies, annotation protocols, acquisition pipelines, and report styles can cause models to entangle clinical semantics with dataset identity, leading to poor transfer despite increased scale. In this work, we revisit radiology VLM construction from the perspective of harmonized multi-source learning. We propose HarMoE, a dataset-aware mixture-of-experts framework that learns shared cross-dataset medical semantics while confining source-specific variation to lightweight residual experts in deeper decoder layers. To further exploit clean supervision from labeled datasets, we train in a unified disease vocabulary with masked multi-dataset supervision, enabling the model to leverage complementary annotations without introducing false negatives. Experiments on large-scale chest X-ray benchmarks show that HarMoE consistently improves zero-shot classification, out-of-distribution transfer, and grounding over strong baselines. Our results suggest that building robust radiology VLMs requires moving beyond single-source image-report alignment toward structured knowledge construction from heterogeneous datasets with cleaner supervision and broader coverage. Code and the 873k harmonized dataset will be released at https://github.com/Roypic/harmoe.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑