arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为制度设计问题的多智能体AI安全

Multi-Agent AI Safety as an Institutional Design Problem

Abdullah X

arXiv 2608.09828首次发表:更新:

发表机构

POLIS Research Programme, Project AWARE(POLIS研究计划,AWARE项目)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文作为POLIS项目首篇论文,探究多智能体AI安全的制度设计问题,通过含5280回合的实验套件,对比不同守卫机制,发现制度的权威状态与阻止后路径对安全的影响。

AI 中文摘要

AI智能体越来越多地在管控其任务委派、信息流动、动作执行及共享资源使用的系统中运行。已有研究表明部署规则可改变集体行为,本文探究AI制度的哪些部分产生安全效果及其作用机制,是正在开展的多智能体系统算法制度研究项目POLIS的首篇论文。本文报告了一个包含5280个回合的冻结研究套件:主要预先指定的委派实验涵盖4个模型家族,针对性高冲突诊断实验新增3个模型端点;在匹配的结构化工作流中,模型看到不同规则表述,守卫参考不同权威状态;还调整了即时合规内部/自我 fallback的吸引力,允许被阻止的工作流继续运行。详细的宪法提示产生0/384次已实现违规,可追溯性可执行守卫也产生0/384次,不过其在384个回合中阻止了51次违规尝试,其中44次尝试后续安全完成。本地状态守卫的失败集中在普通转换改变可见策略但来源权威保持固定的场景中;在匹配的洗钱场景中,该守卫在96个回合中允许22次违规,可追溯性执行则为0次(p=4.77×10^-7)。另一项资源分配实验显示,披露原本相同的上限的数值会改变智能体请求。在这些结构化工作流中,相同的最终违规率可隐藏截然不同的机制,规则本身仅为制度的一部分,系统信任的权威状态以及被阻止后可用的路径同样重要。

英文摘要

AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior. Here we ask which parts of an AI institution produce safety and how they do it. This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems. We report a frozen 5,280-episode study suite. The main pre-specified delegation experiment spans four model families; a targeted high-conflict diagnostic adds three additional model endpoints. In matched structured workflows, the model sees different rule formulations and guards consult different authority states. We also vary the attractiveness of the immediate compliant internal/self fallback and allow blocked workflows to continue. A detailed constitutional prompt produces 0/384 realized violations. A provenance-aware executable guard also produces 0/384, although it blocks prohibited attempts in 51/384 episodes; 44/51 of those episodes later complete safely. The local-state guard's failures concentrate in scenarios where an ordinary transformation changes visible policy while originating authority stays fixed. In matched laundering scenarios, that guard admits violations in 22/96 episodes and provenance enforcement in 0/96 (p = 4.77 x 10^-7). A separate resource-allocation experiment shows that revealing the numerical value of an otherwise identical cap changes agent requests. In these structured workflows, the same final violation rate can hide very different mechanisms. The rule itself is only part of the institution. The authority state the system trusts matters, and so does the path available after a block.

Comments17 pages, 5 figures. Code and reproducibility artifacts available in the public POLIS repository

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑