AI 中文总结
研究智能体安全评估结果缺乏承重证据问题,引入属性级可重构性度量及跨线束适配器,通过特定协议和方法定义相关指标,在公共和捆绑轨迹上验证,为安全评估提供可重构性度量及相关验证方法。
AI 中文摘要
许多智能体安全评估结果并非承重证据:相同的名义结果(任务成功、攻击成功、监控分数)可能基于截然不同的证据体系。目前尚无供应商中立的、可运行的工具来衡量作为评估有效性指标的可重构性,即捕获的证据能否重构相关决策。本文引入了一种针对八个决策属性类别的属性级可重构性度量,以及一个跨线束适配器,它会为每次运行的监控覆盖发布检查生成支持每个决策的证据充分性卡片……
英文摘要
Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes. No vendor-neutral, runnable instrument scores reconstructability as an evaluation-validity metric: whether captured evidence can reconstruct the decision a claim depends on. This paper introduces a property-level reconstructability metric over eight decision-property classes and a cross-harness adapter emitting per-decision Evidence Sufficiency Cards backing a per-run monitor-coverage release check. It specifies a counterfactual-replay intervention protocol, implements its replayability-precondition probe, and defines a claim-evidence overclaim gap. On public and bundled traces, without new model runs, twelve-field sufficiency spans 0.458-0.833 across four inputs sharing a surface reading; replay preconditions are unmet in every scored trace. In a synthetic release-gate pair, the sufficiency gate blocks the raw variant (0.542) and passes the instrumented (0.667). Safety-evaluation claims should travel with their reconstructability vector; a reproducibility package regenerates every reported number.
Comments36 pages, 3 tables. Reproducibility package (scorer, fixtures, Evidence Sufficiency Cards): https://doi.org/10.5281/zenodo.21055696