发表机构
Dhirubhai Ambani University(达鲁巴伊·安巴尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对临床问答,提出基于UMLS概念重叠的软验证器与熵归一化LLM评判器及一致性惩罚的复合奖励,在GRPO中提升模型性能,优于SFT。
AI 中文摘要
语言模型的强化学习后训练依赖于两种奖励设计:人类偏好(RLHF、DPO)和二元验证器(RLVR)。临床问答两者都不适用。接近正确的答案仅因一个替换的实体而不同,且没有可执行的检查来决定临床正确性。我们从维护的受控词汇表中实例化一个软验证器:UMLS概念唯一标识符重叠(通过scispaCy,集合级F1)提供一个分级的、外部指定的奖励,计算时无需模型参与循环。我们将其与GRPO内的熵归一化LLM评判器结合,该评判器覆盖了重叠无法看到的安全性和证据轴,以及一个小的对填充和重复的一致性惩罚,使早期训练样本可评分。这个三项复合奖励在MedQA上相比SFT对Phi-3-mini(3.8B)的EM提高了2.9%(0.700对0.680),Token-F1提高了39%(0.202对0.145);对Llama-3.2-3B的相应提升为EM提高14%,Token-F1提高35%。我们报告Token-F1作为主要指标,因为它认可了EM在此开放生成规模下丢弃的部分正确的临床内容。主表结果是3个种子的平均值,标准差低于0.005。该方法迁移到PubMedQA,在PubMedQA训练集上使用相同的复合奖励训练,无需重新调整参数,Phi-3-mini的Token-F1相比SFT提高了22%,Llama-3.2-3B提高了17%。在Phi-3上的奖励消融,在固定一致性权重下改变评判器-本体拆分,将3个EM点归因于本体项,该贡献捕获了评判器无法捕获的实体替换。三个负面发现限制了设计:随机负样本下的DPO对强先验模型表现不如SFT,但对最弱先验模型有帮助;稀疏神经奖励下的PPO发散;KL损失内的GRPO在7B时崩溃。
英文摘要
Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a soft verifier from a maintained controlled vocabulary: UMLS Concept Unique Identifier overlap (via scispaCy, set-level F1) gives a graded, externally specified reward computed without a model in the loop. We combine it inside GRPO with an entropy-normalised LLM judge, which covers the safety and evidence axes overlap cannot see, and a small consistency penalty on padding and repetition that keeps early-training samples scorable. This three-term composite improves over SFT on Phi-3-mini (3.8B) over MedQA by 2.9% on EM (0.700 vs 0.680) and 39% on Token-F1 (0.202 vs 0.145); on Llama-3.2-3B the corresponding gains are 14% on EM and 35% on Token-F1. We report Token-F1 as the primary metric because it credits partially-correct clinical content that EM discards at this open-generation scale. Main-table results are means over 3 seeds with standard deviations below 0.005. The method transfers to PubMedQA, where training on the PubMedQA train set with the same composite reward improves Token-F1 over SFT by 22% on Phi-3-mini and 17% on Llama-3.2-3B without retuning. A reward ablation on Phi-3, varying the judge-ontology split at a fixed consistency weight, attributes 3 EM points to the ontology term, the contribution that catches entity substitutions the judge cannot. Three negative findings constrain the design: DPO under random negatives underperforms SFT for strong-prior models but helps the weakest-prior one; PPO under a sparse neural reward diverges; GRPO with KL-in-loss collapses at 7B.
CommentsAccepted at CIKM 2026