发表机构
The Chinese University of Hong Kong, Shenzhen; Shenzhen Research Institute of Big Data; Shenzhen Loop Area Institute; National Health Data Institute, Shenzhen(香港中文大学(深圳); 深圳市大数据研究院; 深圳河套学院; 国家健康数据研究院(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对仅RL领域自适应中的梯度饥饿和教师分布锚定问题,提出OnePO方法,通过自适应目标演化与教师退休机制,在医学领域以少量样本超越SFT+RL和纯RL,并构建开源模型HuatuoGPT-3。
AI 中文摘要
领域自适应旨在将通用大语言模型(LLM)转变为目标领域的专家。虽然主流的SFT+RL流水线提供了便捷的冷启动方式,但它可能降低探索多样性,并通过多阶段优化引入额外复杂性。这些局限性促使了仅RL自适应的研究。然而,纯在线策略RL面临冷启动问题,而混合策略RL仍存在不足:教师输出中的信息性标记在早期训练中学习过慢,而过时的教师输出可能阻碍后续改进。我们将这两种失败模式识别为梯度饥饿和教师分布锚定。为解决这些问题,我们提出单阶段策略优化(OnePO),将教师输出视为策略改进的临时指导。OnePO结合自适应目标演化以加强对信息性低概率教师标记的学习,以及教师退休机制,在当前策略能够超越教师输出时将其丢弃。在医学自适应中,OnePO在仅使用20K训练样本的情况下,在HealthBench(总计)上达到67.2分,分别比SFT+RL和纯RL高出2.7分和7.4分。我们进一步扩展OnePO以生成HuatuoGPT-3,这是一个开源医学LLM系列,其27B变体在HealthBench(总计)上达到70.1分,在HealthBench专业上达到71.4分,超越了GPT-6 Astra等前沿模型。模型和代码可在该https URL获取。
英文摘要
Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative tokens in teacher outputs are learned too slowly in early training, and stale teacher outputs can hinder later improvement. We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring. To address them, we propose One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance for policy improvement. OnePO combines Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once the current policy can surpass them. On medical adaptation, OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively. We further scale OnePO to produce HuatuoGPT-3, an open-source medical LLM series whose 27B variant reaches 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional, surpassing frontier models such as GPT-6 Astra. Models and code are available at https://github.com/FreedomIntelligence/HuatuoGPT-3.
CommentsExtended version of "OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation", accepted at ICML 2026, with additional analysis and scaling to HuatuoGPT-3