arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20538cs.AI

拒绝、分解、刷新:面向闭环AI评估的声明安全协议

Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation

Peiying Zhu, Sidi Chang

首次发表
浏览论文内容

中文总结 AI 辅助

针对闭环AI评估中可复现但声明错误的风险,提出拒绝、分解、刷新三动作协议,通过弃权(不执行)、分层报告和参考重算,实现可观测支持与统计校准的可执行契约。

中文摘要 AI 辅助

一项AI评估可以完全可复现,但仍可能支持错误的声明。这一风险在闭环系统中尤为突出:策略决定了访问的状态、可观测的组件以及哪些故障会留下可测量的痕迹。我们提出了一种包含三个动作的声明安全协议。拒绝:当缺乏干净的参考流或匹配的运行时比较支持时,弃权(不执行)。分解:将协议执行、操作性误接纳和结构性假设分别报告,而不是作为一个整体的通过/失败标签。刷新:将分布偏移警报视为使参考映射失效并重新计算的请求,而非故障证据。我们在一个仅聚合的模拟器中实例化了该协议,该模拟器包含24个策略组件、三种需求模式、两个故障掩码族以及独立的开发和保留种子。预注册的保留集包含1,440个案例和21,600个分区行。只有55/72个模式-组件单元被参考接纳,且54/55个保持运行时接纳,这使得弃权(不执行)成为结果的一部分。稳定误接纳为0/20个已表示组件,在冻结的0.20规则下,单侧精确95%上限为0.1391。在被接纳的单元内,受影响的干净流量预测超过了名义故障单元比例:在嵌套于20个组件簇中的540个单元-臂行中,单元减去流量的负对数似然差为每行0.1264纳特,95%组件簇区间为[0.0593, 0.1918]。漂移日志显示了为什么“空”必须是参考相对的:干净的故障-空流在三种模式下触发了15/15、0/15和14/15次警报,而只有中间模式与冻结的检测器参考匹配。我们贡献的不是一个通用阈值,而是一个可执行的契约,将可观测支持、统计校准和合理声明联系起来。

英文摘要

An AI evaluation can be perfectly reproducible and still support the wrong claim. This risk is acute in closed-loop systems: policy determines visited states, observable components, and which failures leave a measurable trace. We propose a claim-safe protocol with three actions. Refuse: abstain when a clean reference stream or matched runtime comparison lacks support. Decompose: report protocol execution, operational false admission, and structural hypotheses separately rather than as one PASS/FAIL label. Refresh: treat distribution-shift alarms as requests to invalidate and recompute a reference map, not as fault evidence. We instantiate the protocol in an aggregate-only simulator with 24 policy components, three demand regimes, two fault-mask families, and independent development and heldout seeds. The preregistered heldout contains 1,440 cases and 21,600 partition rows. Only 55/72 regime-component units were reference-admitted and 54/55 remained runtime-admitted, making abstention part of the result. Stable false admission was 0/20 represented components, with a one-sided exact 95% upper bound of 0.1391 under a frozen 0.20 rule. Within admitted units, affected clean traffic outpredicted nominal fault-cell fraction: across 540 unit-arm rows nested in 20 component clusters, the cell-minus-traffic negative-log-likelihood difference was 0.1264 nats per row, with a 95% component-cluster interval of [0.0593, 0.1918]. A drift log shows why "null" must be reference-relative: clean fault-null streams triggered 15/15, 0/15, and 14/15 alarms across three regimes, while only the middle regime matched the frozen detector reference. Rather than a universal threshold, we contribute an executable contract linking observable support, statistical calibration, and justified claims.

发表机构

  • Blossom AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑