发表机构
Harvard Medical School/MGH; Zhejiang University; Harvard University(哈佛医学院/麻省总医院; 浙江大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对医学多项选择问答中测试时强化学习因答案空间结构而崩溃的问题,提出PROSE方法,以过程奖励模型对推理步骤评分并取最小值作为轨迹奖励,显著提升通用Llama模型性能,超越专用医学模型并匹敌更大系统。
AI 中文摘要
测试时强化学习利用多数投票伪标签在模型自身的无标注测试集上对模型进行适配,并在数学领域展现出强劲结果。我们证明,这一方法在医学多项选择题问答中失效:准确率停滞不前,而输出多样性迅速下降。通过一项受控实验——保持问题、模型和优化器不变,仅改变答案空间——我们将这一失败归因于答案空间的结构,而非领域难度。在较小的答案空间中,错误的采样轨迹常常碰撞到同一个错误的伪标签并强化它;在较大的答案空间中,它们分散开来,获得的奖励很少。这一诊断催生了PROSE(过程奖励引导的自训练),它奖励推理质量而非答案一致性。PROSE使用医学过程奖励模型对每个推理步骤进行评分,将轨迹奖励设为各步骤得分的最小值,并强制执行答案格式约束。在无标签的情况下,PROSE显著改进了一个通用Llama模型,超越了专门构建的医学模型,并匹敌更大的系统。由于过程信号被内化到策略中,适配后的模型在推理时无需奖励模型,并将其增益迁移到未见过的数据集上。我们进一步证明,最小聚合至关重要:均值聚合可能被利用,使代理奖励饱和而准确率下降。
英文摘要
Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathematics. We show that this recipe collapses on medical multiple-choice QA: accuracy stagnates while output diversity rapidly declines. Through a controlled experiment that keeps the questions, model, and optimizer fixed while changing only the answer space, we trace this failure to answer-space structure rather than domain difficulty. In small answer spaces, incorrect rollouts often collide on the same wrong pseudo-label and reinforce it; in large answer spaces, they disperse and receive little reward. This diagnosis motivates PROSE, Process Reward Guided Self-Training, which rewards reasoning quality instead of answer agreement. PROSE scores each reasoning step with a medical process reward model, assigns the trajectory reward as the minimum score across steps, and enforces answer-format constraints. Without labels, PROSE substantially improves a general Llama model, surpassing purpose-built medical models and matching much larger systems. Because the process signal is internalized into the policy, the adapted model requires no reward model at inference and transfers its gains to unseen datasets. We further show that the minimum aggregation is essential: mean aggregation can be exploited, saturating the proxy reward while degrading accuracy.