arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI控制的可靠性理论

Reliability Theory for AI Control

Grant Molnar

arXiv 2609.26419首次发表:更新:

AI 中文总结

本文应用可靠性理论分析AI控制栈的故障抑制能力,提出通过Birnbaum重要性指导组件改进,以提升系统可靠性并应对恶意部署风险。

AI 中文摘要

可靠性理论为分层系统提供了成熟的语言,但其形式化工具尚未成为前沿AI控制的标准。我们将这些工具应用于谷歌DeepMind针对恶意部署的防御措施。同一控制栈根据其故障域的不同,可以实现三次方、二次方或线性的罕见故障抑制。Birnbaum重要性识别出哪些组件改进能带来最大的标称可靠性提升,而预防措施则改变了需要恢复的人群。这些结果为分离、改进、测量和测试的内容提供了具体指导。

英文摘要

Reliability theory gives a mature language for layered systems, but its formal tools are not yet standard in frontier AI control. We apply them to Google DeepMind's defenses against rogue deployment. The same control stack can have cubic, quadratic, or linear rare-failure suppression depending on its failure domains. Birnbaum importance identifies which component improvements buy the most nominal reliability, while prevention changes the population on which recovery is demanded. These results give concrete guidance about what to separate, improve, measure, and test.

Comments14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑