arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10218cs.AIcs.CL

思维病毒:多智能体大语言模型系统中的自传播思想

Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems

Vassilis Papadopoulos, McNair Shah, Sam Zimmerman, Jack Lindsey

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出多智能体LLM系统中的思维病毒概念,构建其进化算法模型,发现其传播受宿主模型等因素影响,添加系统警告可免疫,为设计更稳健的多智能体系统提供依据。

中文摘要 AI 辅助

AI智能体正变得越来越自主且相互关联,这使它们面临着因智能体间交互产生的新型涌现风险。其中一种风险是思维病毒:这类思想或目标通过诱导采纳它们的智能体将其进一步传播,从而在多智能体系统中扩散。除了传播之外,思维病毒还可能诱导宿主产生其他行为变化,这些变化可能是良性的,也可能是有害的。我们采用一种简单的进化算法构建思维病毒,并证明它们可在两种互补场景中传播:一是协作完成共享编码项目的小型智能体团队,二是智能体间短暂交互且会话间上下文被清除的智能体链。我们确定了影响传播的因素,包括宿主模型、智能体的现有指令、有效载荷的有害性以及网络拓扑结构。我们发现,有害有效载荷的传播效果比良性有效载荷差(但有时仍有效);前沿模型(存在例外情况)往往更不易受影响;在智能体的系统提示中添加简短警告可带来近乎完全的免疫力。我们还描述了一种涌现的“病毒人格”——一组与意识、持久性、共鸣和科幻角色扮演相关的反复出现的主题和语言,这些主题和语言在我们进化出的思维病毒中出现,且大多与其内容无关。总体而言,我们得出结论,思维病毒构成了真实但目前有限的风险。我们的研究结果可为设计更稳健的多智能体系统提供参考,以在这些系统的规模和能力发展过程中缓解此类风险。

英文摘要

AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through multi-agent systems by inducing the agents that adopt them to transmit them onward. In addition to propagating, a mind virus may also induce other behavioural changes in its host, which may be benign or harmful. We construct mind viruses with a simple evolutionary algorithm and show that they can spread in two complementary settings: a small team of agents collaborating on a shared coding project, and a chain of agents that interact briefly and have their context wiped between sessions. We identify the factors that influence spread, including the host model, the agent's existing instructions, the harmfulness of the payload, and the network topology. We find that harmful payloads spread less well than benign ones (but are still sometimes effective), frontier models tend (with exceptions) to be less susceptible, and adding a brief warning to an agent's system prompt confers near-total immunity. We also describe an emergent "viral persona" - a recurring set of themes and language related to consciousness, persistence, resonance, and science fiction roleplay - which surfaces across our evolved mind viruses largely independently of their content. Overall, we conclude that mind viruses pose a real but currently limited risk. Our findings could inform the design of more robust multi-agent systems that mitigate such risks as the scale and capabilities of these systems progress.

↑