arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.06596cs.CRcs.LG

校准家族过拟合:为何可信破坏监测器不能跨谱系转移

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages

Lucas Pinto

首次发表
浏览论文内容

中文总结 AI 辅助

研究可信监测器跨谱系转移问题,以不可信策略家族为受控轴,通过拟合和应用监测器分解跨家族AUROC,发现相互作用为正且差距几何,给出四步协议,指出控制评估应报告跨家族转移矩阵,单一配对准确性高估安全性。

中文摘要 AI 辅助

可信监测是人工智能控制中的核心防御手段:一个成本较低的可信模型对不可信模型的行为进行破坏评分,对最可疑的行为进行审核或延迟处理。此类监测器针对一两个不可信模型进行评估,并将准确性作为监测器的一个属性报告。我们探讨这是否部分取决于配对关系。我们将不可信策略家族作为受控轴:在家族A的记录上拟合一个监测器,将其应用于家族B,并将跨家族的AUROC分解为每个家族破坏行为的明显程度、每个监测器的能力以及去除这两者后的剩余家族自身优势:即相互作用。在代码后门记录上,这种相互作用是正向的,并且在主要混淆因素存在的情况下依然存在:在严格无泄漏的基础上,四个开放权重家族上为+0.172(95%置信区间[+0.158, +0.185]),在跨越四个前沿供应商、公然匹配控制、高达34B参数的经过训练的记录读取监测器以及第二个代码基准的八家族矩阵中保持一致。这种差距是几何性的:每个谱系的破坏行为占据不同的评分方向,将监测器的评分方向转向目标谱系可恢复大部分损失的检测能力,但匹配的随机旋转则无法恢复任何检测能力。在1%的审核预算下,在“ defer-to-trusted”策略下,非谱系监测器捕获19%的攻击尝试,而谱系内监测器在相同审核成本下捕获41%,差距为+21.7个百分点。控制评估应报告跨家族转移矩阵:单一配对的准确性高估了监测器在未校准谱系上提供的安全性。我们给出了一个针对这种差距的四步协议,每一步都是一个实测结果。

英文摘要

Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicious are audited or deferred. Such monitors are evaluated against one or two untrusted models, and the accuracy is reported as a property of the monitor. We ask whether it is partly a property of the pairing. We make the untrusted policy family the controlled axis: we fit a monitor on family A's transcripts, apply it to family B, and decompose the cross-family AUROC into how obvious each family's sabotage is, how capable each monitor is, and the residual own-family advantage after both are removed: the interaction. On code-backdoor transcripts the interaction is positive and survives the dominant confounds: +0.172 (95% CI [+0.158, +0.185]) on four open-weight families on a strict leak-free basis, holding across an eight-family matrix spanning four frontier vendors, blatancy-matched controls, a trained transcript-reading monitor up to 34B parameters, and a second code benchmark. The gap is geometric: each lineage's sabotage occupies a different scoring direction, and rotating the monitor's scoring direction toward the target lineage recovers most of the lost detection while a matched random rotation recovers nothing. At a 1% audit budget under defer-to-trusted, an off-lineage monitor catches 19% of attack attempts where an in-lineage monitor catches 41% at the same audit cost, a +21.7-point gap. Control evaluations should report cross-family transfer matrices: a single-pairing accuracy overstates the safety a monitor delivers against a lineage it was not calibrated on. We give a four-step protocol that acts on the gap, with each step a measured result.

发表机构

  • Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑