AI 中文总结
本研究提出亚十亿参数规模的开源医疗语言模型MedLLM,经三阶段流程训练,在医疗基准测试中展现出亚十亿规模特有的能力分化现象,为医疗小模型研究提供了新方向。
AI 中文摘要
开源医疗语言模型已趋于统一规模:所有广泛使用的系统均采用7B参数及以上规模,亚十亿参数领域尚未得到充分研究。我们提出MedLLM,这是一款参数规模为0.1B的开源医疗语言模型,通过完全开源的三阶段流程训练:采用课程序列长度调度的通用预训练、在MedFineWeb上进行的领域微调(MedFineWeb是我们发布的参考引导医疗语料库,通过与医疗问答数据的嵌入相似度从通用网络数据中筛选得到),以及结合监督微调(SFT)与直接偏好优化(DPO)的偏好对齐微调。在医疗基准测试中,MedLLM呈现出仅在亚十亿规模下可见的模式:医疗能力在模型压缩后并非均匀下降,而是按任务类型分化。在基于上下文的问答任务中,MedLLM与适配医疗场景的7B模型仅相差2.9个百分点,且优于指令微调及通用7B基线模型;在知识回忆类问答任务中,其在临床 vignette 类MedQA任务中接近任务下限,但在MedMCQA任务中显著超越所有7B及亚7B基线模型,表明在回忆能力不足时,限制因素是模型容量而非适配性。这种能力分化在7B规模下被掩盖(该规模下两种能力均具备),仅在容量稀缺时才会显现。
英文摘要
Open medical language models have converged on a single scale: every widely used system runs at 7B parameters or more, leaving the sub-billion regime uncharacterized. We present MedLLM, an open 0.1B-parameter medical language model trained through a fully open three-phase pipeline: general pretraining with curriculum sequence-length scheduling, domain fine-tuning on MedFineWeb, a reference-guided medical corpus we release that is selected from general web data by embedding similarity to medical question-answering (QA) data, and preference-aligned fine-tuning combining SFT with direct preference optimization (DPO). Across medical benchmarks, MedLLM shows a pattern visible only at sub-billion scale: medical competence does not degrade uniformly under compression but splits by task type. On context-grounded QA it comes within $2.9$pp of a medically adapted 7B model and surpasses the instruction-tuned and general-purpose 7B baselines; on knowledge-recall QA it stays near the task floor on clinical-vignette MedQA yet significantly exceeds every 7B and sub-7B baseline on MedMCQA, indicating that where recall fails the constraint is model capacity rather than adaptation. This dissociation is masked at 7B, where both capabilities are present, and surfaces only when capacity is scarce.