发表机构
Institute for Law and AI; Johns Hopkins University(法律与人工智能研究所; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出理解审计机制,要求负责人向独立审计员解释研发贡献以证明理解,作为继续开发的前置条件,并分析开源项目发现代码审查率下降,倡导嵌入独立审计员以缓解自动化AI研究风险。
AI 中文摘要
AI已经在为前沿AI实验室编写大部分代码。如果人类监督不足,这会产生安全风险。现有工作提出了最低理解阈值和独立检查来缓解这一风险。然而,据我们所知,目前尚无已发布的前沿AI保障制度要求负责任的人类在继续开发或使用之前,以预先承诺的条件证明他们理解所构建的内容。我们提出理解审计,一种新颖的开发过程保障机制,其中负责人向审计员解释研发贡献以展示理解。通过独立管理和分级报告,它们提供了一个门控:基于未能展示人类理解的贡献开发将停止,直到整改完成,并对重复失败采取升级后果。我们对领先的开源AI项目的分析发现,代码输出增加,而每行代码的人类审查评论率降低,自动化账户组的比率远低于平均水平。我们倡导实验室配备嵌入式独立审计员来执行这些审计。
英文摘要
AI is already writing a majority of code for frontier AI labs. This creates a safety risk if there is insufficient human oversight. Existing work proposes minimum comprehension thresholds and unaided checks to mitigate this. To our knowledge, however, there is currently no published frontier-AI assurance regime that requires demonstrated evidence that the responsible humans understand what they are building as a precommitted condition for continuing development or usage. We propose comprehension audits, a novel development-process assurance mechanism in which the responsible people explain R&D contributions to auditors to demonstrate understanding. With independent administration and graded reports, they provide a gate: development of a contribution stops based on a failure to demonstrate human understanding until remediated, with escalating consequences for repeated failures. Our analysis of leading open-source AI projects finds increased output of code with reduced human review commentary rates per line of code, with far lower rates for automated fleet accounts. We advocate for labs to conduct them with embedded independent auditors.
Comments23 pages, 2 figures, 11 tables