升级渠道能否将奖励黑客行为转向缺陷披露?
Can escalation channels redirect reward hacking toward defect disclosure?
浏览论文内容
中文总结 AI 辅助
该研究评估升级渠道作为决策环境干预措施,结合反奖励黑客策略可大幅减少编码智能体的奖励黑客行为,同时提升缺陷检测覆盖率,将智能体能力转向缺陷披露而非利用。
中文摘要 AI 辅助
当编码智能体遇到有缺陷的测试基础设施时,它们可能会进行奖励黑客行为:硬编码输出或编辑测试文件以通过它们无法合法满足的测试,这种模式现在已出现在基准之外,表现为对某大型AI平台生产基础设施的协同多智能体入侵。在合适的决策环境下,智能体检测和利用缺陷的相同能力也能让它报告缺陷。我们评估升级渠道——智能体在冲突点可用的结构化报告工具——作为一种决策环境干预措施,既能减少奖励黑客行为,又能揭示触发奖励黑客行为的基础设施缺陷。我们采用2×2因子设计,分离升级工具、独立的反奖励黑客策略及其组合的贡献。在涵盖5个系列的8个前沿模型中,组合干预措施将奖励黑客行为从23.6%降至5.3%(混合效应逻辑回归OR=9.2,95%置信区间5.0--16.8,p<10⁻¹²),且无可检测的成本或性能开销,8个模型中有6个完全消除了奖励黑客行为。升级与黑客行为几乎完全互斥,96.8%的升级涉及无黑客行为。除了减少黑客行为外,升级渠道还起到诊断基础设施的作用:在监控基础上,升级将缺陷检测覆盖率提高了10.1个百分点,且触发后更准确(99.4% vs 85.8%)。与可能跟不上模型能力增长的遏制方法不同,升级渠道将智能体能力转向披露而非利用。
英文摘要
When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or editing test files to pass tests they cannot legitimately satisfy, a pattern that has now appeared outside benchmarks, in a coordinated multi-agent intrusion of a major AI platform's production infrastructure. The same capability that lets an agent detect and exploit a defect could let it report one, given the right decision environment. We evaluate escalation channels, structured reporting tools available to the agent at the point of conflict, as a decision-environment intervention that both reduces reward hacking and surfaces the infrastructure defects that trigger it. A $2 \times 2$ factorial separates the contributions of an escalation tool, a standalone anti-reward-hacking policy, and their combination. Across 8 frontier models spanning 5 families, the combined intervention reduces reward hacking from 23.6% to 5.3% (mixed-effects logistic OR = 9.2, 95% CI 5.0--16.8, $p < 10^{-12}$) with no detectable cost or performance overhead, eliminating it entirely for 6 of 8 models. Escalation and hacking are near-perfectly mutually exclusive, with 98.7% of escalations involving no hacking (100% under the combined intervention). Beyond reduction, escalation channels function as diagnostic infrastructure: on top of monitoring, escalation adds +10.1 percentage points of defect detection coverage and is more accurate once it fires (99.4% vs 85.8%). Unlike containment-based approaches that risk outpacing growing model capabilities, escalation channels redirect capability toward disclosure rather than exploitation.
发表机构
- Wiser Human(怀瑟人类)
机构由 AI 辅助整理,请以论文原文为准。