arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

奖励黑客攻击下的评估器集成:协方差几何与有限搜索保证

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

Fariya Afrin, Ibne Farabi Shihab

arXiv 2608.08002首次发表:更新:

发表机构

Kalinga Institute of Industrial Technology; Iowa State University(卡林加工业技术学院; 爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对奖励黑客攻击下评估器集成的失效问题,基于协方差几何分析,提出有限搜索的保证方法,通过压力测试与模型审计验证理论,明确分歧诊断的局限。

AI 中文摘要

语言模型评判器和奖励模型支持可扩展监督,但有限优化可能会利用评估器错误而非提升响应质量。我们通过评估器集成的协方差几何来表征这种失效:对于校准后的评判器,集成均值会沿全1方向保留共模误差,而评判器间分歧仅捕获正交误差,因此分歧可能在聚合稳健时仍很高,或在存在共享响应依赖误差时仍很低。我们证明仅从内部评判器分数无法识别共模误差;在联合次高斯模型下,我们对K选最优选择的夸大程度和目标质量遗憾进行了界定,将保证扩展至条件校准下可预测的自适应搜索,所得搜索项随log K的平方根缩放,对高斯投影误差是渐近紧的。我们进一步表明,有噪质量代理会引入人工秩1协方差且不改变分歧,并提出一种有界双锚点伯恩斯坦证书用于有限搜索误差与遗憾。针对120组(J, ρ, K)配置的固定种子高斯压力测试及真实模型审计验证了该理论,同时揭示了在搜索压力增大时,基于分歧的诊断方法的局限性。

英文摘要

Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this failure through the covariance geometry of evaluator ensembles. For calibrated judges, the ensemble mean retains common-mode error along the all-ones direction, whereas cross-judge disagreement captures only orthogonal error. Consequently, disagreement can be high despite robust aggregation, or low while shared response-dependent errors persist. We prove that common-mode error is not identifiable from internal judge scores alone. Under a joint sub-Gaussian model, we bound best-of-K selection overstatement and target-quality regret, extending the guarantees to predictably adaptive search under conditional calibration. The resulting search terms scale as the square root of log K and are asymptotically tight for Gaussian projected errors. We further show that noisy quality proxies introduce artificial rank-one covariance without changing disagreement, and propose a bounded two-anchor Bernstein certificate for finite-search error and regret. Fixed-seed Gaussian stress tests over 120 (J, rho, K) configurations and real-model audits validate the theory while revealing the limits of disagreement-based diagnostics under increasing search pressure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑