arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越正确性:面向生物医学大语言模型评判者的有效性导向评估

Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges

Rodrigo de Oliveira, Federico Pittino, James Gwinnutt, Jay Nanavati

arXiv 2608.29127首次发表:更新:

发表机构

IQVIA(艾昆纬)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对生物医学LLM评判者,该研究提出有效性导向的评估流程,评估Llama-3.1-8B-Instruct的四种训练场景,发现SFT→RL场景在多维度表现最优,可匹配或超越前沿模型。

AI 中文摘要

在高质量人工评判稀缺时,我们提出一种可扩展、面向有效性的流程,用于评估生物医学大语言模型(LLM)评判者。首先,我们通过基于确定性指标的变异扩充现有带人工标注的生物医学基准,生成可审计的偏好对。其次,我们从三个与部署相关的维度评估评判者,而非仅看整体正确性:与指标衍生的黄金标签的正确性、重复随机采样下的鲁棒性、对请求输出格式的合规性。我们用该流程在四种场景下评估Llama-3.1-8B-Instruct:(1)基础场景,直接使用指令调优模型;(2)SFT场景,仅基于蒸馏的监督微调;(3)RL场景,仅基于GRPO的强化学习;(4)SFT→RL场景,先SFT后RL。基础场景和单阶段场景在结构化医学判别(如PICO提取、临床计算)上表现不佳,而SFT→RL在正确性、合规性和鲁棒性上表现最佳;其提升集中于可分解任务(PICO、MedCalc),有时可匹配或超越前沿模型。

英文摘要

We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond aggregate correctness using three deployment-relevant dimensions: correctness against metric-derived gold labels, robustness under repeated stochastic sampling, and compliance with the requested output format. We use this pipeline to assess Llama-3.1-8B-Instruct under four regimes: (1) base, using the instruct model as is; (2) SFT, distillation-based supervised fine-tuning only; (3) RL, GRPO-based reinforcement learning only; and (4) SFT$\rightarrow$RL, SFT followed by RL. The base and single-stage regimes struggle on structured medical discrimination such as PICO extraction and clinical calculations, whereas SFT$\rightarrow$RL performs best across correctness, compliance, and robustness; gains concentrate on decomposable tasks (PICO, MedCalc), at times matching or outperforming frontier models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑