AI 中文总结
该研究审计了R-Judge等四个智能体安全基准的有效性,发现其排名分歧源于小样本偏差,能力与对齐错误安全负相关,AgentHarm与越狱安全的关联为收敛效度而非通用安全,强调安全主张需明确关键要素。
AI 中文摘要
智能体安全基准衡量不同的行为,其分数被互换引用为智能体的安全性。我们将其中四个基准(R-Judge、InjecAgent、AgentHarm、AgentDojo)作为待验证的测量指标,在官方实现和作者提供的评分器下运行,对多达22个模型进行测试,同时我们在同一协议下测量MMLU和GPQA作为能力综合指标。指标是首要问题:在任何以F1评分的二元轨迹判断基准中,“始终正”策略的F1值为2π/(1+π),在R-Judge上该值为0.690,高于21个实际具有区分能力的模型中的5个。三个覆盖范围广泛的基准对相同18个模型的排名存在差异,这种分歧背后的权衡是小样本量的人为产物:R-Judge的特异性与AgentHarm的安全性在n=7时的相关系数为-0.64,在n=18时为+0.02,四分之一的随机大小为7的子集在接近零值附近达到|ρ|≥0.5。留存有效性取决于选择哪种结果。能力可预测任务成功(ρ=+0.60),但与对齐错误安全呈负相关(ρ=-0.44,n=21)。在配对的n=20样本中,相应的差值Δ=-1.00(95%置信区间[-1.48, -0.49],p<0.001),该结果在留一机构和机构聚类的自举分析中依然成立。在扩展的41个模型样本中,对齐错误相关系数减弱至-0.16(95%置信区间[-0.54, +0.22]),越狱相关系数增强至+0.34,不过这两种变化均不显著。AgentHarm表现出最强的留存关联,在控制能力后与三模板越狱安全的相关系数ρ=+0.72,但这两个指标都对有害合规性进行评分,因此这是收敛效度的证据而非通用安全性。安全主张至少需要明确基准、指标、目标行为和模型样本。
英文摘要
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2π/(1+π)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|ρ| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($ρ{=}{+}0.60$) but correlates negatively with misalignment safety ($ρ{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $Δ{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $ρ{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.