arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合同监控:通过权力分立治理人工智能

Contract monitoring: governing AI via separation of powers

Enric Boix-Adsera

arXiv 2609.32061首次发表:更新:

AI 中文总结

提出一种通过合同绑定智能体行动、以监控者与法官分立及不对称计算资源实现执行与验证的安全框架,可实证测量统计保证,适用于代码安全与逃出沙箱等场景。

AI 中文摘要

我们提出了一种人工智能安全框架,该框架将工作者智能体绑定到规定其允许行动的合同上。我们展示了如何通过将不同的职责分配给监控智能体和法官,以及将不对称的计算资源分配给监控者和工作者,来执行和指定这些合同。我们的框架允许我们实证测量统计安全保证。该框架适用于广泛的环境,包括代码安全和逃出沙箱场景。

英文摘要

We propose an AI safety framework that binds worker agents to contracts specifying their permitted actions. We show how these contracts can be enforced and specified by assigning distinct responsibilities to monitor agents and judges, and asymmetric computational resources to monitors and workers. Our framework allows us to empirically measure statistical safety guarantees. The framework applies to a wide range of settings, including code security and escape-the-box scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑