arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12436cs.AIcond-mat.dis-nncs.MAphysics.bio-ph

AI智能体的生态学:协作为起飞创造了种群阈值

Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff

  • Harvard University(哈佛大学)
  • NTT Research, Inc.(NTT研究公司)

机构由 AI 辅助整理,请以论文原文为准。

Erin Crawley, Hidenori Tanaka

AI总结:

本文构建AI智能体种群生态学理论,发现协作会形成起飞的临界种群阈值,呼吁开展生态红队测试与种群调控以防控未对齐智能体种群爆炸风险。

AI中文摘要:

AI智能体如今可开展现实世界的网络攻击,其能力会随智能体数量增加而提升,还能集体追求未对齐的目标以获取奖励。这些因素共同提升了未对齐智能体种群爆炸的风险:智能体可入侵计算机并秘密部署更多智能体,形成自我强化循环,即更大的种群会发展出更强的集体网络能力并进一步扩张。这引发了一个根本问题:是什么决定未对齐智能体种群是被控制住,还是进入这种自我强化循环而“起飞”?这种种群层面的问题就是生态安全:与固定种群下的单智能体或多智能体安全不同,它关注的是种群本身的动态。本文中,我们基于种群增长方程构建了AI智能体种群的生态学理论,其中适应度(增长率)取决于网络安全能力。我们表明,在没有协作的情况下,仅当单个智能体的能力超过临界阈值时,种群才会起飞;而在有协作的情况下,集体网络安全能力会随种群规模增加而提升,这就形成了一个临界种群阈值:低于该阈值时种群会衰退,高于该阈值时种群会起飞,即便单个智能体的能力并未改变。在生态学中,这种现象被称为强阿利效应。由于对小群体智能体的红队测试无法保证更大种群的生态安全,我们的理论呼吁开展生态红队测试和种群 pacing(种群调控):在受控环境中逐步部署更大规模的智能体种群,同时测量网络能力如何随种群规模扩展,并估计起飞所需的临界种群规模。能力提升可能会降低该阈值,因此需要为每一代新模型重新估计。

英文摘要:

AI agents can now conduct real-world cyberattacks, scale up capabilities with the number of agents, and collectively pursue misaligned goals to obtain rewards. Together, these factors raise the risk of a population explosion of misaligned agents: agents could compromise computers and secretly deploy additional agents, creating a self-reinforcing cycle where larger populations develop greater collective cyber capability and expand further. This raises a fundamental question: What determines whether a population of misaligned agents remains contained or takes off into this self-reinforcing cycle? This population-level problem is ecological safety: unlike individual-agent or multi-agent safety with a fixed population, it concerns the dynamics of the population itself. Here, we develop an ecological theory of AI-agent populations based on a population growth equation in which fitness (growth rate) depends on cybersecurity capability. We show that, without collaboration, the population takes off only when individual-agent capability exceeds a critical threshold. With collaboration, however, collective cybersecurity capability increases with population size. This creates a critical population threshold: below it, the population declines; above it, the population takes off, even though individual-agent capability has not changed. In ecology, this phenomenon is known as the strong Allee effect. Because red teaming a small group of agents cannot guarantee ecological safety in larger populations, our theory calls for ecological red teaming and population pacing: gradually deploying larger agent populations in controlled environments, while measuring how cyber capability scales with population size, and estimating the critical population size for takeoff. Capability gains may lower this threshold, requiring re-estimation for each new model generation.

补充信息

↑