arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

旗帜游戏:一种用于机制性群体可解释性的玩具模型

Flag Game: A Toy Model for Mechanistic Swarm Interpretability

Elizabeth Pavlova, Hidenori Tanaka

arXiv 2609.19124首次发表:更新:

发表机构

Harvard University; NTT Research, Inc.; Cambridge Boston Alignment Initiative(哈佛大学; NTT研究公司; 剑桥波士顿对齐计划)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出旗帜游戏玩具模型,研究集体信念形成机制,发现群体规模导致信念崩溃转向极化,并引入社会电路归因与统计力学理论解释,迈向机制性群体可解释性。

AI 中文摘要

AI智能体的涌现协调行为正开始带来严重的安全风险。驱动这些行为的一个关键现象是关于世界的信念的快速形成与传播,而机制性理解对于集体对齐至关重要。为此,我们引入了旗帜游戏,一个用于研究集体信念形成机制的玩具模型。具体而言,一个隐藏的国家旗帜定义了真实标签,每个有界智能体仅直接观察到一个私有裁剪区域,但可以交换信念并权衡来自同伴的社会证据。尽管其简单,旗帜游戏再现了丰富的集体现象:性能随群体规模的非单调缩放、来自社会意识提示和团队多样性的准确性提升,以及组织结构的强烈影响。特别是,我们识别出,在较小群体规模下的集体信念崩溃随着群体增长转变为集体信念极化。这种极化导致了在较大群体规模下的性能下降,但创造了集体信念的多样性。最后,我们通过两种互补方法剖析了集体信念崩溃和极化背后的机制。我们首先引入社会电路归因,一种预测哪个智能体以及什么观点对集体动态最重要的技术,并通过因果干预智能体来验证其预测,追踪智能体修补如何改变集体结果。然而,对智能体的因果干预的有效性随着群体增长而降低。因此,我们为较大群体开发了一种统计力学理论,并验证其与经验相图匹配。总之,这些结果向机制性群体可解释性迈出了第一步,这是一门关于个体智能体的属性及其通信如何产生涌现集体行为的科学。

英文摘要

Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for collective alignment. To this end, we introduce the Flag Game, a toy model for studying the mechanisms of collective belief formation. Concretely, a hidden country flag defines the ground truth, and each bounded agent directly observes only a private crop but can exchange beliefs and weigh social evidence from peers. Despite its simplicity, the Flag Game reproduces rich collective phenomenology: non-monotonic scaling of performance with population size, accuracy gains from social-awareness prompting and team diversity, and strong effects of organizational structure. In particular, we identify that collective belief collapse at small population sizes turns into collective belief polarization as the population grows. This polarization causes the performance decline at large population sizes, but creates diversity in collective beliefs. Finally, we dissect the mechanisms underlying collective belief collapse and polarization with two complementary approaches. We first introduce social circuit attribution, a technique to predict which agent, and what view, matters most to collective dynamics, and verify its predictions by causal interventions on agents, tracing how agent patching changes collective outcomes. However, the efficacy of causal interventions on agents decreases as the population grows. We therefore develop a statistical mechanical theory for larger populations and verify that it matches the empirical phase diagram. Together, these results take a first step toward mechanistic swarm interpretability, a science of how the properties of individual agents and their communication give rise to emergent collective behavior.

Comments21 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑