监督缺口:LLM 安全监控器遗漏了什么,以及为何这不是能力问题
The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出用全变差度量将LLM安全监控的不可判定性转化为可检测性前沿,发现监控器缺陷主要源于信息与程序缺失而非能力不足,并强调超属性基准需机械验证。
AI中文摘要:
安全监控器被要求证明的若干属性,包括跨租户非干扰、消极怠工(sandbagging)和评估感知(evaluation awareness),都是 2-安全超属性(2-safety hyperproperties),只能通过两次执行来见证。标准推论是一个二元不可能性:单条轨迹无法判定这些属性。我们将二元性替换为一种度量。一个紧致界将任何单轨迹监控器的平衡准确率限定为 $\ frac12+\ frac12\,TV(P_0,P_1)$,从而将不可判定性转化为一个分级可检测性前沿,并定义了一个监督缺口:监控器相对于该前沿的不足。在一个具有闭式全变差(TV)的泄漏族上,九个 LLM 监控器在 $TV=0$ 时达到最优,但随着 $TV$ 增长仅捕获少量信号;在 $TV=1$ 时,一个 20 行的成员检查得分达到 100%,而它们平均仅为 60.9%。这种不足主要不是能力问题:明确指出要检查的内容弥补了 61% 的缺口,同时将 $TV=0$ 的对照保持在随机水平。同样的区分贯穿于一个 $2\ imes2$ 因子设计:一个想象的第二次运行使监控器停留在随机水平(50.4%),而在实际执行的第二次运行上应用相同规则则达到 90.0%,并且一个存储的预言机(oracle)若没有比较程序则仅产生 68.2% 的结果。信息和程序各自都是必要的,且两者都不是能力。在非确定性下,重放仅在正确的投影下才能跟踪闭式 $k$-重放曲线,并且一个投影前沿表明由此产生的困境是不可避免的:窄投影漏掉 98.6% 的带外泄漏,宽投影标记 75.7% 的干净流量,而可达到的准确率随良性变化率和通道数量按 $1/(qm)$ 衰减。最后,两个前沿 LLM 评审员认证了我们自己基准的早期版本为健全的,而一个符号检验发现了一个方向性偏差($p=2.7\ imes10^{-5}$),该偏差使我们的三项发现无效。超属性基准的构造有效性应通过机械方式证明,而非由模型审计。
英文摘要:
Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hyperproperties, witnessed only by two executions. The standard consequence is a binary impossibility: one trace cannot decide them. We replace the binary with a measurement. A tight bound puts the balanced accuracy of any single-trace monitor at $\tfrac12+\tfrac12\,TV(P_0,P_1)$, turning undecidability into a graded detectability frontier and defining an oversight gap: a monitor's shortfall below it. On a leak family with closed-form $TV$, nine LLM monitors are optimal at $TV=0$ but capture little signal as $TV$ grows; at $TV=1$, where a 20-line membership check scores $100\%$, they average $60.9\%$. That shortfall is mostly not capability: naming what to check closes $61\%$ of it while leaving the $TV=0$ control at chance. The same split runs through a $2{\times}2$ factorial: an imagined second run leaves monitors at chance ($50.4\%$) while the same rule on an executed second run reaches $90.0\%$, and a stored oracle without a comparison procedure yields only $68.2\%$. Information and procedure are each necessary and neither is capability. Under nondeterminism, replay tracks a closed-form $k$-replay curve only under the right projection, and a projection frontier shows the resulting dilemma is forced: narrow misses $98.6\%$ of off-channel leaks, broad flags $75.7\%$ of clean traffic, and attainable accuracy decays like $1/(qm)$ in the benign-variation rate and the channel count. Finally, two frontier LLM judges certified an earlier version of our own benchmark as sound while a sign test found a directional bias ($p=2.7\times10^{-5}$) that invalidated three of our findings. Construction validity for hyperproperty benchmarks should be proved mechanically, not audited by models.