arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能红队评估能证明什么以及不能证明什么

What AI Red-Team Evaluations Can and Cannot Prove

Bandana Kaur

arXiv 2607.21735首次发表:更新:

发表机构

APIsec Research Labs(APIsec研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人工智能红队评估能证与不能证之事,通过定义证据上限确定界限,发现高于某危害率基准可证类别,低于则否,该界限不限于基准,审核评估套件发现当前基准对高频危害足够,对罕见灾难性不足。

AI 中文摘要

人工智能模型的红队评估支持一些主张而不支持其他主张,两者之间的界限是可计算的,而非仅仅是判断问题。我们将评估的证据上限定义为在固定测试预算下一个结果能改变信念的最大因素,以封闭形式推导基准零结果的证据上限,并用其精确确定界限。我们发现,高于可计算的危害率时,适度规模的基准能以规定的证据标准证明一个类别,此时无故障结果比单个重现的失败结果更强。低于该率时,在固定评分规则和近似独立试验结构下,可行规模的被动基准无法提供指定的安全证据,两种情况的交叉点有封闭形式。该界限不限于基准,还涵盖自适应和自动化红队评估,表明假设间的区分而非攻击成功决定证据价值。通过对照界限审核八个评估套件,我们发现当前基准对高频危害类别足够,但对罕见灾难性类别短几个数量级。安全基准并非无信息,它们对特定且可计算的命题集有信息,关键是要说明是哪些命题。

英文摘要

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.

Comments21 pages, 4 figures, 5 tables. Code and data links provided in the manuscript. v2: corrected Figure 1(b); corrected required sample sizes in Table 4 and in Sections 4.2, 4.6 and 5.2, which had been rounded rather than taken to the ceiling; corrected the sample-size expression stated in Methods; minor corrections to Table 1 and the Figure 2 caption. No theorem, result or conclusion is affected

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑