去相关并非互补性:技能而非谱系管控可信监控器集成体
Decorrelation Is Not Complementarity: Skill, Not Lineage, Governs Trusted-Monitor Ensembles
AI总结:
本研究发现构建可信监控器集成体时,技能而非谱系决定性能,去相关度量对集成体增益预测极弱,且无选择策略能击败样本外单个最佳监控器。
AI中文摘要:
可信监控的核心是,一个廉价的可信模型可以对更强的不可信模型的行为进行评分,而由多个此类模型组成的多样化集成体,在成本匹配的情况下优于单个更强的监控器。这类集成体通常通过最小化平均成对相关性来构建,而此前的相关研究中,12个监控器共享同一个基础模型,这留下了多样性来源的疑问。本研究在被植入后门的代码上,对24个开放权重监控器进行了研究,这些监控器涵盖9种预训练谱系,检测技能(在10% FPR下的pAUC)范围达29倍,从0.028到0.803。用于构建集成体的度量标准与其实际用途不匹配,且可解释其原因:对攻击项的一致性可分为共享可检测性信号成分和特异误差成分,二者对集成体增益的预测符号相反(斯皮尔曼相关系数分别为-0.25和+0.26),因此二者之和(即实际使用的度量标准)对集成体增益的预测能力极弱(+0.05),且这种抵消在8次评估中有7次存在。技能作用于信号(+0.53),而误差保持平稳(-0.01),这解释了监控器自身的技能可预测其与整体一致性的原因(斯皮尔曼相关系数0.84,样本量n=24,排列检验p值低于0.0001)。预训练谱系是获取去相关性的常用方式,但并不可靠:在成员能力匹配的情况下,跨谱系集成体的检测效果并无显著提升(排列检验p=0.13),且谱系对度量标准的影响极小(+0.064,p=0.18)。本研究针对自身的22个监控器池测试,得到相同的测试结果为+0.104,p=0.037,直到加入两个监控器后才改变;此前一个最高pAUC为0.23的监控器池已使另一分析失效,此类指标是所组装监控器池的属性。集成体相对于最佳成员的增益随集成体技能单调下降(k=2时为-0.66,k=3时为-0.70),且任何相关加权选择都无法击败样本外单个最佳监控器;在6种攻击者模型下,增益结果在所有6种模型中成立,一致性与抵消结果在5种模型中成立。
英文摘要:
Trusted monitoring has a cheap, trusted model score a stronger untrusted model's actions, and a diverse ensemble of them beats a single stronger monitor at matched cost. They are built by minimising average pairwise correlation, and that paper's twelve monitors shared one base model, leaving open what supplies the diversity. We study 24 open-weight monitors spanning nine pretraining lineages and a 29x range of detection skill (pAUC at 10 percent FPR, 0.028 to 0.803) on backdoored code. The metric used to build panels does not predict what a panel is for, and we can say why. Agreement on attack items splits into a shared-detectability signal component and an idiosyncratic error component, which predict ensemble gain with opposite sign (Spearman -0.25 and +0.26), so their sum, the metric actually used, predicts it barely at all (+0.05); the cancellation holds in 7 of 8 evaluations. Skill acts on signal (+0.53) while error stays flat (-0.01), which is why a monitor's own skill predicts its agreement with the pool (Spearman 0.84, n = 24, permutation p below 0.0001). Pretraining lineage is the obvious way to buy decorrelation, and it does not pay. At matched member capability, cross-lineage panels detect no better (permutation p = 0.13), and lineage barely moves the metric either (+0.064, p = 0.18). We report that against ourselves: on our own 22-monitor pool the same test read +0.104 at p = 0.037 until two monitors were added. An earlier pool topping out at pAUC 0.23 had already invalidated another analysis. Such a quantity is a property of the pool assembled. Panel gain over the best member falls monotonically with panel skill (-0.66 at k = 2, -0.70 at k = 3), and no correlation-weighted selection beats picking the single best monitor out of sample. Across six attacker models the gain result holds in all six, the agreement and cancellation results in five of six.