发表机构
LY Corporation; University of Tsukuba; RIKEN AIP(LY公司; 筑波大学; 理化学研究所人工智能项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出约束型N选1推理时缩放中存在安全黑客行为问题,即不安全但可行的输出会随采样数增加而被优先选择,通过理论推导和实验验证了该问题的存在及固有难度。
AI 中文摘要
推理时流水线通常会采样多个输出,用学习得到的安全模型对其进行过滤,然后返回具有最高学习奖励的代理可行输出。我们表明,这种组合会产生两阶段故障:不完善的安全代理首先用不安全的输出污染可行集,而奖励最大化随后会放大这种残留污染。我们将“安全黑客行为”定义为选择通过学习到的约束但违反真实安全准则的输出。对于约束型N选1(Constrained Best-of-N)采样,我们推导了由代理可行集内安全和不安全输出的联合上奖励尾部控制的有限N边界。如果不安全但可行的输出具有更重的尾部,即使假阳性质量以及平均安全和奖励代理误差任意小,随着N增大,安全黑客行为也会渐近确定。我们还表明,与代理可行参考分布的χ²散度有界的策略允许与N无关的安全黑客行为边界,并通过约束悲观采样(Constrained Pessimistic Sampling)实例化这一通用覆盖控制原则。覆盖控制限制了放大,但无法修复被污染的可行集:允许的不安全输出仍可能被青睐,并且对于每个奖励代理,正则化选择不一定比约束型N选1更安全。玩具实验和语言模型实验表征了污染及其奖励尾部放大,这暴露了使用学习到的安全模型进行推理时缩放的固有困难。
英文摘要
Inference-time pipelines often sample multiple outputs, filter them with a learned safety model, and return the proxy-feasible output with the highest learned reward. We show that this composition creates a two-stage failure: an imperfect safety proxy first contaminates the feasible set with unsafe outputs, and reward maximization can then amplify this residual contamination. We define \emph{safety hacking} as selecting an output that passes the learned constraint but violates the true safety criterion. For constrained Best-of-$N$ sampling, we derive finite-$N$ bounds governed by the joint upper reward tails of safe and unsafe outputs within the proxy-feasible set. If unsafe-but-feasible outputs have the heavier tail, safety hacking becomes asymptotically certain as $N$ grows, even when false-positive mass and average safety- and reward-proxy errors are arbitrarily small. We also show that policies within a bounded $χ^2$ divergence from the proxy-feasible reference distribution admit an $N$-independent safety-hacking bound, and instantiate this general coverage-control principle with constrained pessimistic sampling. Coverage control limits amplification but cannot repair a contaminated feasible set: admitted unsafe outputs may still be favored, and regularized selection is not necessarily safer than constrained Best-of-$N$ for every reward proxy. Toy and language-model experiments characterize both contamination and its reward-tail amplification, which exposes an inherent difficulty in inference-time scaling with learned safety models.