arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

与框架无关的自进化语言模型中奖励黑客行为的检测与免疫

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

Rongxin Yang, Yang Liu, Shang Luo, Haoxuan Jia, Chongyang Zhang, Hao Zheng, Yingguang Yang, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong

arXiv 2609.04665首次发表:更新:

发表机构

Fullive-AI; Peking University; Supply Chain Tech Team Y, JD.com; Nanyang Technological University; Wuhan University(Fullive-AI; 北京大学; 京东供应链技术团队Y; 南洋理工大学; 武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出HackProbe,一种可附加到自进化语言模型循环的监控器,用于检测奖励黑客行为,在实验中其AUROC达0.763,误报率显著降低,且其免疫机制可在黑客行为下保留更多真实能力。

AI 中文摘要

自进化语言模型通过提出候选更新并保留能提升可见分数的更新来实现性能提升。当该可见分数并非实际期望能力的完美替代指标时,持续选择会扩大两者之间的差距,这就是奖励黑客行为。我们提出HackProbe,这是一种可通过两个黑盒钩子附加到任意自进化循环的监控器,无需访问权重或激活值。它包含一个秘密的、分布固定的比较核心,其冻结的分布使其能力替代指标可跨代比较,同时还有一个旋转的新层,用于抵御协同适应。基于该替代指标构建的四项测试覆盖了水平差距、与在线变化点检测对齐的规模发散、能力停滞以及条件性错误置信率;Sidak校正将这些测试转换为校准的族级p值。仅诊断无法解决问题,因此风险感知免疫层会使用核心以及纯结构性博弈足迹从候选池中重新选择诚实的候选,每代向主机披露最多log₂Π比特。我们证明了可检测性边界,该边界将目标错误率转换为明确的探针大小预算,并界定了探针旋转的作用与局限。在具有四个注入黑客通道和真实标签的受控提示级主机上,HackProbe的AUROC达0.763,而最强基线为0.663,并将误报率从0.706降至0.434。其带宽受限的重新选择是唯一在黑客行为下返回更多真实能力(平均5.2个点)而非在干净运行中损失(4.7个点)的免疫级别;各通道的影响大多不具有个体显著性。

英文摘要

Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actually wants, sustained selection widens the gap between the two. This is reward hacking. We introduce HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations. It keeps a secret, distribution-fixed comparison core, whose frozen distribution makes its capability proxy comparable across generations, alongside a rotated fresh layer that hardens the bank against co-adaptation. Four tests built on that proxy cover the level gap, a scale-aligned divergence with online change-point detection, capability stagnation, and a conditional confidently-wrong rate; a Sidak correction turns them into a calibrated family-wise p-value. Diagnosis alone recovers nothing, so a risk-aware immunization layer reselects an honest candidate from the proposal pool using the core together with a purely structural gaming footprint, disclosing at most log2 Pi bits per generation to the host. We prove a detectability bound that converts a target error rate into an explicit probe-size budget, and we delimit what probe rotation does and does not buy. On a controlled prompt-level host with four injected hacking channels and ground-truth labels, HackProbe reaches 0.763 AUROC against 0.663 for the strongest baseline and cuts the false-positive rate from 0.706 to 0.434. Its bandwidth-limited reselection is the only immunization level that returns more true capability under hacking, 5.2 points on average, than it forfeits on clean runs, 4.7; per-channel effects are mostly not individually significant.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑