发表机构
The Hong Kong Polytechnic University; PolyU-Daya Bay Technology and Innovation Research Institute(香港理工大学; 香港理工大学大亚湾技术创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InfiMed2提出阶段感知数据设计与稳定性监督,构建4B和27B医学多模态模型,在五个基准上以66.73%和73.72%的平均准确率超越更大模型。
AI 中文摘要
近期医学多模态模型受益于更大的语料库、更广泛的模态覆盖以及更强的推理导向训练,然而在持续预训练(CPT)和后训练阶段进行有效的数据设计仍然具有挑战性。医学数据源在结构、粒度和信息密度上差异显著,并且随着训练从广泛知识获取推进到后期巩固,其效用也会发生变化。同时,后训练阶段常以短问答形式为主,对信息丰富且答案一致的推理过程提供的监督有限。我们提出InfiMed2,一个包含4B和27B参数的通用医学多模态基础模型系列,其核心是阶段感知的数据设计。我们构建了一个包含55.68B token的语料库,通过特定来源的处理,将广泛的临床知识与富含上下文的生物医学视觉证据相结合。我们的CPT流程首先适配视觉编码器,然后构建广泛的医学知识,最后在学习率衰减阶段过渡到以证据为中心的数据混合。在监督微调(SFT)中,我们使用答案稳定性、答案掩码重建和正确性约束选择来重新生成视觉问答响应,以提供更具信息量和答案一致性的监督。4B模型进一步通过可验证奖励的强化学习(RLVR)进行优化。在五个医学多模态基准测试中,InfiMed2-4B在RLVR后平均准确率达到66.73%,超过了更大的Qwen3.5-9B,而InfiMed2-27B达到73.72%,在评估的开源权重模型中最高。
英文摘要
Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information density, and their utility shifts as training progresses from broad knowledge acquisition to late-stage consolidation. Meanwhile, post-training is often dominated by short-form visual question answering, providing limited supervision for informative and answer-consistent explanations. We introduce InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design. We curate a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing. Our CPT pipeline first adapts the vision encoder, then builds broad medical knowledge, and finally transitions to an evidence-focused data mixture during learning-rate decay. For supervised fine-tuning (SFT), we regenerate visual question-answering responses using answer stability, answer-masked reconstruction, and correctness-constrained selection to produce more informative and answer-consistent supervision. The 4B model is further optimized with reinforcement learning with verifiable rewards (RLVR). Across five medical multimodal benchmarks, InfiMed2-4B achieves 66.73% mean accuracy after RLVR, surpassing the larger Qwen3.5-9B, while InfiMed2-27B reaches 73.72%, the highest among the evaluated open-weight models.