静默修订:衡量前沿AI开发者安全框架中未披露的变更
Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出静默修订率指标,通过分析十二家前沿AI开发者安全框架的版本语料,发现多数实质性变更未被披露,且削弱承诺的变更更常被静默处理,主张发布义务应附带枚举义务。
AI中文摘要:
前沿AI开发者发布安全框架,承诺证明其模型是否危险。欧盟和加利福尼亚州现在将这些文件视为问责工具,并且两者都已对其修订施加义务。两者都不要求修订是可读的,即读者能够从开发者自己的叙述中了解发生了什么变化。我们引入了静默修订率,即开发者发布的叙述未识别的框架承诺实质性变更所占的比例,并发布了用于计算该指标的分版本、哈希固定的语料库。该语料库包含十二家已发布安全框架的开发者所有公开版本,以及每家提供方的变更日志、红线标注或公告。我们追踪了十二个连续版本对中的710个承诺实例,根据冻结的编码手册进行编码,并逐一裁决了244个实例。由此得出三项发现。第一,在严格标准下,67%的实质性变更(95%置信区间62至72)是静默的,在宽松标准下为53%,在章节粒度上降至49%。第二,静默性似乎与叙述形式相关,因为叙事性公告的静默率为74%,而逐项变更日志为63%,而叙述长度(以字数计)几乎无关紧要;在尊重嵌套结构的测试中,差异具有提示性。第三,77%的被追踪变更削弱或移除了承诺,并且在八对中有七对中,削弱比加强更常是静默的。因此,法定补救措施存在,但指定了错误的工件。理由说明解释了框架为何变更,枚举说明陈述了变更内容,只有后者才能使修订可审计。我们认为,发布义务应附带枚举义务,而一家提供方已经自愿且不完整地满足了这一义务。
英文摘要:
Frontier AI developers publish safety frameworks that commit them to evidencing whether their models are dangerous. The European Union and California now treat these documents as instruments of accountability, and both already impose duties on their revision. Neither requires the revision to be legible, in the sense that a reader could learn from the developer's own account what changed. We introduce the silent revision rate, the share of material changes to a framework's commitments that the developer's published account does not identify, and we release the versioned, hash-pinned corpus needed to compute it. The corpus contains every public version of the safety frameworks of the twelve developers that have published one, together with each provider's changelog, redline or announcement. We trace 710 commitment instances across twelve consecutive version pairs, code them against a frozen codebook, and adjudicate 244 individually. Three findings follow. First, 67% of material changes (95% CI 62 to 72) are silent under a strict standard and 53% under a lenient one, falling to 49% at section granularity. Second, silence appears to track the form of the account, since narrative announcements run at 74% against 63% for itemised changelogs, whereas account length in words barely matters; on the test that respects nesting the difference is suggestive. Third, 77% of traced changes weaken or remove a commitment, and in seven of eight pairs weakenings are more often silent than strengthenings. The statutory remedy therefore exists and specifies the wrong artefact. A justification explains why a framework changed, an enumeration states what changed, and only the latter makes revision auditable. We argue that publication duties should carry an enumeration duty, which one provider already meets, voluntarily and incompletely.