关于AI安全评估的翻译注释
A Translational Note on AI Safety Evaluation
浏览论文内容
中文总结 AI 辅助
本文指出AI安全评估中自动化红队优于人工的结论存在衡量与声称错位,提出“威胁模型覆盖缺口”问题,并论证需部署情境不同的评估者来弥补。
中文摘要 AI 辅助
近期研究表明,在标准AI安全基准测试中,自动化红队测试比人工红队测试能以更低成本发现更多漏洞,一些人将此视为人类评估者正变得可有可无的证据。然而,该比较衡量的是某一事物,而结论声称的却是另一事物。基准测试衡量的是攻击者在开发者预先设定的、固定的危害集合中搜索的彻底程度,而未被纳入该集合的危害对于任何在该集合内工作的攻击者(无论自动化与否)都是不可见的。同样的盲点也出现在学术密码学和临床药物试验中,在这些领域,内部有效的评估对其从未指向的人群保持沉默。我们将AI安全领域的这一现象称为“威胁模型覆盖缺口”,并发现它在当前的开源权重模型中持续存在,其中危害出现在英语基准测试遗漏的非英语提示中。弥补这一缺口需要部署情境与开发者不同的评估者。支持这些评估者的理由是方法论层面的,基于覆盖范围,而现有的评估框架不太可能自行产生这样的评估者。
英文摘要
Recent studies report that automated red-teaming finds more vulnerabilities, at lower cost, than human red-teaming on standard AI safety benchmarks, and some read this as evidence that human evaluators are becoming dispensable. The comparison measures one thing and the conclusion claims another. A benchmark measures how thoroughly an attacker searches a predefined set of harms, fixed in advance by the developers, and a harm left out of that set is invisible to any attacker working inside it, automated or not. The same blind spot appeared in academic cryptography and in clinical drug trials, where an evaluation that was internally valid stayed silent about the population it was never pointed at. We call the AI-safety version the \emph{threat-model coverage gap}, and find that it persists in a current open-weight model, where harms surface in non-English prompts that English benchmarks miss. Closing it requires evaluators whose deployment context differs from the developers'. The case for those evaluators is methodological, grounded in coverage, and the existing evaluation frame is unlikely to produce them on its own.