arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信任不是分数:高风险AI智能体的运行时保障契约

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

Serhii Zabolotnii

arXiv 2609.39717首次发表:更新:

发表机构

Cherkasy State Business College; State Scientific Research Institute of Armament and Military Equipment Testing and Certification; healthPrecision(切尔卡瑟国立商学院; 国家武器装备与军事装备试验认证科学研究所; healthPrecision)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高风险AI智能体在关键任务中证据与权限脱节的问题,提出运行时保障契约(RAC),通过强制门禁和证据记录约束行动,经故障注入实验验证其优于纯分数规则。

AI 中文摘要

基准测试、审计和智能体协议描述了性能、权限和修复,但并未说明在关键任务过程中,观察到的证据应如何改变智能体的权限。我们将此称为保障过渡缺口。我们提出运行时保障契约(RAC),这是一种策略级形式化模式,将自主边界、组件资格、证据状态、过渡策略、人工审查能力和不可补偿门禁绑定在一起。在RAC下,软指标可指导路由,而失败或未知的强制门禁则强制重试、切换、升级、推迟或停止;聚合性能不能授权行动。我们定义了契约、证据记录、权限规则和五个不变量,并在临床、工业和司法失败探针中进行了说明。随后,我们报告了一项在智能体编码中的确定性故障注入研究:280个构造案例,由门禁合取规则、仅分数规则和受限协议基线进行评估。在已发布的示例权重和阈值下,分数规则放行了100个需阻断注入中的80个以及全部40个需审查注入。事后调整后,它在该语料库上与合取规则匹配。对于正权重、正阈值、二元风险信号、零信号对照以及每个信号单独触发的注入案例,我们证明精确一致性成立当且仅当阈值不超过最小权重。另外一组18条手工编写的轨迹检查了版本固定的证据和审查过渡与更简单策略变体的对比。在进一步的前瞻性合成保留集(24个情节)中,两位盲法LLM评审员对所有72次行动尝试分配了相同标签;RAC和单独实现的完整有状态基线均匹配这些标签。这些研究在合成案例上测试了机制;它们既未证明部署安全性,也未证明跨领域有效性。

英文摘要

Benchmarks, audits, and agent protocols describe performance, permissions, and repair, but not how observed evidence should change an agent's authority during a consequential task. We call this the assurance-transition gap. We propose a Runtime Assurance Contract (RAC), a policy-level formal schema binding autonomy boundaries, component eligibility, evidence state, transition policy, human-review capacity, and non-compensatory gates. Under RAC, soft metrics may inform routing, whereas a failed or unknown mandatory gate forces retry, switch, escalation, deferral, or stop; aggregate performance cannot authorize action. We define the contract, an evidence record, a permission rule, and five invariants, and illustrate them in clinical, industrial, and judicial failure probes. We then report a deterministic failure-injection study in agentic coding: 280 constructed cases evaluated by a gate conjunction, a score-only rule, and a restricted protocol baseline. At the published example weights and threshold, the score rule admits 80 of 100 block-required injections and all 40 review-required injections. Tuned in hindsight, it matches the conjunction on this corpus. For positive weights, a positive threshold, binary risk signals, zero-signal controls, and an injected case firing each signal alone, we show that exact agreement holds if and only if the threshold does not exceed the smallest weight. A separate set of 18 hand-authored traces checks version-pinned evidence and review transitions against simpler policy variants. In a further prospective synthetic holdout of 24 episodes, two blinded LLM judges assign identical labels to all 72 action attempts; RAC and a separately implemented full stateful baseline both match these labels. These studies test mechanisms on synthetic cases; they establish neither deployed safety nor cross-domain effectiveness.

Comments16 pages, 2 figures, 5 tables. Ancillary files: decision log, executable transition model, LLM-labelled synthetic holdout. Synthetic mechanism study; no deployment claim

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑