更多弃权(不执行),而非更强判别性:预执行大语言模型监督中的验证单元
More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对预执行LLM监督的验证单元问题,提出双前缀框架,发现一或两个行动的短验证单元信息性最高,更长窗口会使监控器更倾向于弃权而非提升判别性。
AI中文摘要:
预执行监督是人工智能控制中可信监控的核心:一个易出错的大语言模型(LLM)监控器会在不可逆转的执行前审查计划的行动。过度阻止会丧失实用性,并迫使部署者将其禁用。每个协议都必须确定一个验证单元:每次调用审查多少个行动。现有设计将该单元视为给定值;其对易出错监控器的影响尚未被测量。自然轨迹无法将其分离:审查长度与错误类型和位置共同变化。仅以捕获率为指标会产生误导:拒绝一切能捕获所有内容。测量这一点仅需边界变化和匹配的干净对照组。我们引入双前缀框架,该框架同时提供这两者。每个黄金计划产生一个注入了一个环境可接受错误的前缀,以及一个在一次写入上不同的干净双前缀。在五个嵌套长度下判断每对,将裁决变化与该单元单独关联。判别性通过预先注册的信息性(捕获率减去误拒率)评分。更长的审查会提高捕获率;误拒率同步上升。在两个领域的所有六名评判者中,信息性在一或两个行动时达到峰值:更长的窗口使零样本监控器更倾向于弃权(不执行),而非更具判别性。重放保留的观察结果表明,失败主要源于观察缺失。安全论证应说明该单元并共同报告干净序列。我们的框架是第一个针对此选择的受控、预先注册的工具,且绝不单独读取捕获率。我们校准的短单元在八行动审查上恢复了高达0.95的信息性,且没有任何经过测试的无标签政策能持续胜过它。
英文摘要:
Pre-execution oversight is core to trusted monitoring in AI control: a fallible LLM monitor vets planned actions before irreversible execution. Over-blocking forfeits usefulness and pressures deployers to disable it. Every protocol must fix a unit of verification: how many actions one call reviews. Existing designs take the unit as given; its effect on fallible monitors is unmeasured. Natural traces cannot isolate it: review length co-varies with error type and position. Catch alone misleads: rejecting everything catches everything. Measuring this needs boundary variation alone and a matched clean control. We introduce the twin-prefix framework, which supplies both. Each gold plan yields a prefix with one injected, environment-accepted error and a clean twin differing in one write. Judging each pair at five nested lengths ties verdict changes to the unit alone. Discrimination is scored by pre-registered informedness, catch minus false rejection. Longer review raises catch; false rejection climbs in lockstep. Informedness peaks at one or two actions for all six judges in both domains: longer windows make zero-shot monitors more rejective, not more discriminative. Replaying withheld observations traces the failure largely to observation deprivation. Safety cases should state the unit and co-report the clean series. Our framework is the first controlled, pre-registered instrument for this choice and never reads catch alone. Our calibrated short unit recovers up to 0.95 informedness over eight-action review, and no tested label-blind policy consistently beats it.