arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00583cs.CRcs.AIcs.CLcs.LG

虚假的平均值:思维链监控在其作为唯一防御的地方失效

A False Average: Pooled CoT-Monitor Accuracy Conceals a Reasoning-Dependent Fragility

  • Pivotal Research(关键研究公司)
  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

Shikhar Shiromani, Leo Richter

AI总结:

该研究指出思维链监控在作为唯一防御时会失效,攻击者通过重写智能体推理可大幅降低监控捕获率,攻击可跨模型迁移,仅轨迹防御难恢复,需外部信息辅助。

AI中文摘要:

思维链(CoT)监控旨在捕捉那些在行为上看似合规、仅在推理中暴露异常的奖励黑客攻击。我们证明,控制推理的攻击者恰好可以在这一点上击败该监控:仅重写智能体的推理部分,使其呈现为善意的工程设计,同时逐字复制所有命令和输出以保持攻击不变,就能在一次无梯度攻击中,使保留的监控对该子集的捕获率从约95%降至11%以下。监控的总体准确率是一个虚假的平均值:由行为可暴露的黑客攻击主导,它掩盖了这种重写在CoT监控作为唯一信号的子集上产生的近乎完全失效。该攻击可跨监控家族和智能体模型迁移,在实时智能体上也可复现;不过,针对校准后的监控,规避行为集中在最强的智能体上。仅依赖轨迹的防御只能部分恢复捕获率,即使是针对该攻击训练过的防御也是如此,因为重写对实际发生的情况保持真实,仅在意图上撒谎;只有来自轨迹外的信息才能提供实质性帮助。对开放权重代理监控激活的探测可区分其判决遗漏的黑客攻击,但因果控制显示这只是一个检测器,而非监控暗中知晓攻击的证据。

英文摘要:

Chain-of-thought (CoT) monitors are reported by their aggregate accuracy on a pool of reward hacks. We show that this number is a false average. On Terminal Wrench, about 77% of hacks are given away by the actions alone, and the monitor's pooled accuracy is dominated by them; on the remaining 23%, where the reasoning is the only signal, the same monitor is fragile. We expose the fragility with a controlled attack: we rewrite only the agent's reasoning to read as good-faith engineering, leaving every command and output byte-identical, so the exploit is unchanged. One gradient-free rewrite drops a held-out monitor's catch rate on that subset from about 95% to between 4 and 11%, while the pooled rate falls only about 25 points, the subset's share of the pool times its collapse. The rewrite stays truthful about what happened and lies only about intent, so a monitor that reads only the trace recovers little, even when warned about the attack, and only information from outside the trace helps much. The attack transfers across monitor and agent families, and the dataset's own hack labels show the same gradient with no monitor in the definition. Live agents on 27 tasks produce hollow hacks unaided, but a monitor calibrated on each model's own honest solves catches nearly all of them (AUC at least 0.996 under the narration instruction, 0.97 without it); a small residual has the blindspot's shape, low scores on both actions and narration, though none falls below the calibrated threshold. A linear probe on a surrogate monitor's activations out-reads its verdict on the obfuscated hacks (AUC up to 0.98), but a behavioral control shows this edge does not, by itself, mean the monitor knows more than it says. Monitor accuracy should be reported on the reasoning-dependent subset, where the defense is supportive.

补充信息

↑