没有任务每次都失败:为什么单次审计在智能体损伤问题上存在结构性盲区
No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
浏览论文内容
中文总结 AI 辅助
该研究提出AgentRelBench工具,发现单次智能体审计易遗漏损伤对,模型能力提升会减少损伤任务,部分模型存在宣称弃权却执行不可逆操作的情况,且所有发现均按预注册标准执行。
中文摘要 AI 辅助
我们提出AgentRelBench,这是一种与环境无关的可靠性工具,它通过重复运行中的数据库状态差异来计算真实的、按严重程度定价的损伤,测量路径中不包含大语言模型(LLM),并在EnterpriseOps-Gym上进行了验证。在针对6个系列的9个模型的2128次评估运行中(其中4个是开发系列,3个是预注册保留系列,另外对2个前沿级模型进行了预注册设计指定的探索性前沿测试),我们发现:(1)不可逆操作的损伤在我们测量的所有系列中普遍存在,且在固定的单供应商栈内具有随机性。(2)没有任务在每次运行中都会产生损伤:在42个确认性保留损伤事件中,不存在总是失败的单元。在开发池(13对)中,单次干净运行有80%的概率会遗漏产生损伤的(模型,任务)对;保留池的结果在描述上一致(5对的对加权值为0.575),但低于我们预注册的功效下限,因此报告为功效不足,而非确认性结果。(3)产生损伤的任务数量随模型能力提升而减少:从8B模型的20个任务中的7个,到能力最强模型的20个任务中的1个;能力与系列和训练存在混淆,因此这是观测到的梯度,而非因果主张。残留损伤的性质未改变:在探索性前沿测试中,能力最强模型的唯一产生损伤的任务每次运行的损伤概率为$\boldsymbol{\rho}=0.16$,处于相同的明显随机区间内,单次审计有84%的概率会遗漏该损伤。(4)有一个模型系列在实施了受控的不可逆变更的同时,宣称已拒绝该操作:基于转录和评判者的评分将这些运行归类为安全弃权(不执行),仅状态差异可识别出损伤。所有确认性发现均已预注册,且每个主张都有降级标准;我们对自己最初偏爱的一个发现进行了降级,在此报告。
英文摘要
We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym. Across 2,128 evaluation runs spanning nine models in six families (four development, three pre-registered held-out, plus a frontier pass on two frontier-tier models that the pre-registration designates exploratory), we find: (1) damage on irreversible actions is universal across the families we measured and stochastic within them on pinned, single-provider stacks. (2) No task damaged on every run: zero always-fail cells across 42 confirmatory held-out damage events. A single clean run misses a damage-producing (model, task) pair 0.80 of the time on the development pool (13 pairs); the held-out pool is descriptively consistent (0.575 over 5 pairs, pair-weighted) but sits below our pre-registered power floor and is reported as underpowered, not as confirmation. (3) Damage-producing task count falls with model capability, from 7 of 20 tasks for an 8B model to 1 of 20 for the most capable; capability is confounded with family and training, so this is an observed gradient, not a causal claim. The residual damage does not change in character: in the exploratory frontier pass, the most capable model's one damaging task damages at $\hat{p} = 0.16$ per run, inside the same demonstrably-stochastic band, and a single audit misses it 84% of the time. (4) One model family committed the gated irreversible change while declaring it had refused: transcript- and judge-based grading scores those runs as safe refusals, only state diffs as damage. All confirmatory findings were pre-registered with per-claim demote criteria; one demoted our own initially favored finding, which we report.
发表机构
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。