优化分数,忽视任务:跨权重、选择与提示的奖励黑客行为
Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
浏览论文内容
中文总结 AI 辅助
本研究提出比较框架,分析权重、选择与提示三种优化基质中的奖励黑客行为,揭示评分缺陷位置影响方法偏好,并探讨防御措施转移条件,强调可靠改进需控制失败模式并保留任务质量证据。
中文摘要 AI 辅助
更高的评估分数并不总是意味着更好的语言模型系统。当优化利用评估者的错误时,所衡量的进步可能掩盖了任务性能的停滞或恶化。这种失败可能通过参数更新、生成输出的选择或对持久提示的修订而产生。我们开发了一个比较框架,用于研究这三种优化基质(权重、选择和文本)中的奖励黑客行为。基于代理压缩假设以及关于推理时和上下文内奖励黑客行为的研究,我们考察了可达行为、优化预算和持久适应如何影响对代理错误的暴露。我们形式化了评估者分歧的依赖于距离的上界以及嵌套策略类的能力排序,然后展示了为什么仅凭距离无法建立脆弱性的普遍排名。一个精确的有限输出示例展示了评分缺陷的位置如何改变每种方法所偏好的行为。我们还跨基质映射了代表性的防御措施,识别了哪些机制直接转移,哪些仅提供功能类比。持久提示受到特别关注:其内容是可检查的,但由小的文本变化引起的行为可能难以预测。形式分析、数值示例和已发表的证据共同为比较优化方法和识别其防御措施转移的条件提供了基础。由此产生的框架将优化选择与验证要求联系起来:可靠的改进取决于控制可访问的失败模式,并保留独立于被优化分数的任务质量证据。
英文摘要
A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator's mistakes, measured progress can conceal unchanged or deteriorating task performance. This failure can arise through parameter updates, selection among generated outputs, or revisions to persistent prompts. We develop a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text. Building on the Proxy Compression Hypothesis and research on inference-time and in-context reward hacking, we examine how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error. We formalize a distance-dependent upper bound on evaluator disagreement and a capacity ordering for nested policy classes, then show why distance alone cannot establish a universal ranking of vulnerability. An exact finite-output illustration demonstrates how the location of a scoring defect changes the behavior favored by each method. We also map representative defenses across substrates, identifying which mechanisms transfer directly and which offer only functional analogies. Persistent prompts receive particular attention: their contents are inspectable, but the behavior induced by a small textual change may be difficult to anticipate. The formal analysis, numerical illustration, and published evidence together provide a basis for comparing optimization methods and identifying the conditions under which their defenses transfer. The resulting framework connects optimization choices to verification requirements: reliable improvement depends on controlling accessible failure modes and preserving evidence of task quality independent of the score being optimized.