发表机构
University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对智能体社会中通信漏洞导致协作失败的问题,提出在个人约束装置之外增设分层社会约束装置架构,以预防故障、检测无效消息并支持追责。
AI 中文摘要
智能体社会是由多个AI智能体组成的集合,这些智能体代表不同委托方(其目标可能仅部分一致)在信任边界之外自主协调。我们通过实验表明,在智能体社会中,即使诚实且能干的智能体,在现有约束装置和消息原语下也常常无法达成令人满意的结果;而有缺陷或恶意的智能体则可能通过利用通信(“言语”)中的漏洞来阻碍协作、影响结果并追求其他有害目标。我们认为,智能体社会除了每个智能体的“个人约束装置”(用于管理其私有上下文及与其委托方的通信)之外,还需要一个用于智能体间交互的“社会约束装置”。我们提出了一种社会约束装置的分层架构,该架构(i)直接防止某些类别的故障,(ii)使智能体能够在运行时检测无效消息,以及(iii)支持事后调查和追责,并指出了未来研究以实现这些能力的方向。
英文摘要
An agentic society is a collection of AI agents that coordinate autonomously across trust boundaries, on behalf of different principals whose objectives may only partially align. We show experimentally that in agentic societies even honest, competent agents often fail to reach satisfactory outcomes with existing harnesses and messaging primitives, and that faulty or malicious agents can stall collaboration, influence outcomes, and pursue other harmful goals by exploiting vulnerabilities in communication (``speech''). We argue that agentic societies need a \emph{social harness} for inter-agent interactions, in addition to each agent's \emph{personal harness}, which manages its private context and communication with its principal. We propose a layered architecture for social harnesses which (i) prevents classes of failures outright, (ii) enables agents to detect invalid messages at runtime, and (iii) supports post-facto investigation and consequences, and highlight directions for future research to realize these capabilities.