廉价验证器足矣:LLM 后训练对错误奖励具有鲁棒性
A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards
查看机构详情
- ETH Zurich(苏黎世联邦理工学院)
- Handshake AI(Handshake AI公司/Handshake AI(根据上下文推测为公司))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文通过大规模实验证明,在 LLM 后训练中,廉价验证器可达到与昂贵验证器相当的效果,且验证器一致性并非选择最佳训练验证器的可靠指标。
中文摘要 AI 辅助
在具有半可验证奖励的任务上对大型语言模型进行后训练时,从业者必须应对多种因素(训练步数、基础模型大小、训练顺序、数据质量、验证器准确性等)以最大化模型性能。然而,验证器一致性如何预测此类任务上的后训练性能仍不清楚。在本文中,我们通过超过 11,000 个 H100 GPU 小时,在 HealthBench 和 PRBench 任务(涵盖医疗、法律和金融领域)上探索了这一问题。在测试的领域中,Qwen3 受训模型(HealthBench 上为 1.7B-8B;PRBench 上为 8B)、评估划分和前沿 LLM 参考裁判(我们称之为黄金验证器)中,更高的验证器一致性并不能始终识别出最佳的训练验证器。昂贵的验证器不一定优于廉价的验证器,开源的 Gemma 验证器产生了强大的训练结果。我们回顾性地比较了两种低成本选择——一种降低成本的选择和一种平衡的选择——相对于黄金评分协议,估计评分成本降低了 98.8%-99.7%,平均后训练分数与最佳评估训练验证器相差 1-3 分。这些平均值包括个别设置中的较大损失;它们并不表明验证器选择是可互换的。
英文摘要
When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifier agreement predicts post-training performance on such tasks. In this paper, we explore this question with over 11k H100 GPU-hours, across HealthBench and PRBench tasks in medical, legal, and finance domains. Across the tested domains, Qwen3 trainees (1.7B-8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes. We compare two low-cost choices retrospectively -- a cost-reducing choice and a balanced choice -- with estimated grading cost reductions of 98.8%-99.7% relative to the golden grading protocols and average post-training score gaps of 1-3 points from the best evaluated training verifier. These averages include larger losses in individual settings; they do not establish that verifier choices are interchangeable.