已部署验证器套件中的硬门候选资格
Hard-Gate Candidacy in a Deployed Validator Suite
查看机构详情
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究评估部署生成式智能体中13个验证器的硬门候选资格,发现多数检查区分能力弱且探针跳过限制检测上限,强调评估记录需区分未运行检查并携带拒绝证据。
中文摘要 AI 辅助
在验证器被提升为部署流水线上的硬门之前,必须证明其触发能够区分正常到达用户手中的输出与未正常到达的输出。我们在一个已部署的生成式智能体中,对13个验证器进行了该筛选,针对550个运行时构建和350个静态构建(按下游结果标注),并报告每个检查的边际分离度 $J=\mathrm{TPR}-\mathrm{FPR}$,附带Newcombe区间和Fisher精确检验。两个检查在多重比较校正后仍然显著,另外两个仅为名义显著,其余九个与零无法区分,其中三个是因为它们从未在任何采样构建上触发。执行本身相对于被门控的属性并非随机,且这一现象可复现:在覆盖1,867个构建和十个不同运行时检查的四次运行中,探针在895个损坏构建中的144个和972个可接受构建中的1个被跳过(每次运行率为15.6%至16.6%,而可接受构建最多为0.3%),每次跳过都带有相同的“不安全探针”原因。由于被跳过的检查被记录为通过,这施加了一个任何检查质量都无法提升的上限:需要实时工件的检查在此测试平台中操作上最多只能检测约84%的损坏构建。对于唯一具有构造特定标签的检查,一个为空白输出构建的检测器在90个人工标注的空白构建中触发了0个(敏感度的95%上限为3.3%),而其近似的全局帧统计量仅弱区分类别(AUC 0.59),因此该差距并非需要调整的阈值。同样的差距在上一层也出现:在对数万个法官评分构建的普查中,32.5%的拒绝根本没有记录任何问题。我们认为,评估记录必须区分已运行并通过的检查与未运行的检查,必须携带拒绝的证据,并且检查清单并不能作为门控有效性的证据。
英文摘要
Before a validator can be promoted to a hard gate on a deployment pipeline, it has to be shown that its firing separates outputs that reach users in working order from those that do not. We run that screen on 13 validators in a deployed generative agent, against 550 runtime and 350 static builds labelled by downstream outcome, and report each check's marginal separation $J=\mathrm{TPR}-\mathrm{FPR}$ with Newcombe intervals and Fisher exact tests. Two checks survive correction for multiple comparisons, two more are nominal only, and the remaining nine are not distinguishable from zero, three of them because they never fired on any sampled build. Execution itself is not random with respect to the property being gated, and this replicates: across four runs covering 1,867 builds and ten distinct runtime checks, probes were skipped on 144 of 895 broken builds and 1 of 972 acceptable builds (per-run rates 15.6% to 16.6% against at most 0.3%), every skip carrying the same unsafe-to-probe reason. Because a skipped check is recorded as a pass, this imposes a ceiling that no check quality can lift: a check that needs a live artifact cannot operationally detect more than about 84% of broken builds in this harness. For the one check with construct-specific labels, a detector built for blank output fires on 0 of 90 human-labelled blank builds (95% upper bound on sensitivity 3.3%), and the global frame statistic it approximates separates the classes only weakly (AUC 0.59), so the gap is not a threshold that needs tuning. The same gap appears one layer up: on a census of tens of thousands of judge-scored builds, 32.5% of rejections carry no recorded issue at all. We argue that evaluation records must distinguish a check that ran and passed from one that did not run, must carry the evidence for a rejection, and that an inventory of checks is not evidence about a gate.