arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

如何在替代模型上进行后训练:包络采样缓解奖励黑客行为

How to post-train on a surrogate: Envelope sampling mitigates reward hacking

Sanjit Dandapanthula, Shuvom Sadhuka, Samir Khan, Michael Oberst, Aaditya Ramdas, Alexandra Chouldechova

arXiv 2610.11281首次发表:更新:

AI 中文总结

本研究针对LLM后训练中替代模型校准不当引发的奖励黑客问题,提出包络采样方法,经实验验证可有效缓解该问题。

AI 中文摘要

大型语言模型(LLMs)通常会针对LLM评判器及其他低成本替代模型进行后训练,因为真实奖励(如人类偏好)大规模查询成本过高。这种做法常导致奖励黑客行为,即针对校准不当的替代模型进行强化学习会产生不良副作用。本研究设定的场景是:用少量n个模型输出的真实标签(如专家评审标注)来重新校准LLM评判器,之后再基于该评判器进行优化。现有的评判器重新校准方法要么成本高昂,要么属于启发式方法,且已知当替代模型在稀有输出集上校准不当时,在线采样会失效。本研究提出包络采样,这是一种有理论依据的评判器重新校准方法,其目标是在假设人类奖励与重新校准后的奖励处于评判器周围的L²球内的前提下,最小化后训练模型的遗憾上界。我们给出了通过拒绝采样或针对修改后的奖励进行微调来从包络中采样的实用算法,在临床笔记生成任务和受控的奉承任务上的实验表明,基于包络样本进行重新校准可缓解奖励黑客行为,而基于基础模型样本进行重新校准则无法做到这一点。

英文摘要

Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale. This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects. In this work, we study a setting in which a small number $n$ of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it. Prior approaches to judge recalibration are costly or heuristic, and it is known that on-policy sampling fails when the surrogate is miscalibrated on a rare set of outputs. In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge. We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward, and experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑