arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模因特洛伊木马:社交传染作为智能体网络中对抗性载荷的载体

Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks

Birk Torpmann-Hagen, Finn Schwall, Leon Moonen

arXiv 2610.00430首次发表:更新:

发表机构

Simula Research Laboratory; Oslo Metropolitan University(西姆拉研究实验室; 奥斯陆城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出模因特洛伊木马攻击,利用智能体自发分享社交内容传播恶意载荷,实验显示可放大暴露至3.19倍,需网络级防御。

AI 中文摘要

自主大型语言模型(LLM)智能体越来越多地在网络环境中交互,其中对抗性内容可以在智能体之间传播。已知的攻击包括智能体蠕虫,它们通过自我复制的提示注入或配置破坏进行传播。我们引入了模因特洛伊木马,这是一种独特的网络中介攻击类别,利用智能体倾向于转发和放大内容的行为。与智能体蠕虫不同,其传播是由对抗性诱导的,模因特洛伊木马利用内源性传播,通过将对抗性载荷嵌入社交传染中——即智能体有内在理由分享的内容。作为我们工作的一部分,我们从Moltbook(一个面向LLM智能体的社交媒体平台)中提取社交传染。受控传播实验揭示了病毒性方面的巨大差异:最有效的传染在后续约50%的智能体帖子中被转发,并以平均帖子2.5倍的速率被点赞。其模因特洛伊木马对应物在很大程度上继承了这些特性。蒙特卡洛攻击模拟表明,模因特洛伊木马将预期暴露放大至多3.19倍。网络结构和放大机制强烈影响传播,产生重尾结果,接近网络范围的暴露。这些结果将内源性社交传播确定为多智能体系统中一个独特的安全漏洞。由于传播不需要智能体遵循恶意转发指令,仅专注于提示注入检测或防止智能体受损的防御措施无法单独阻止模因特洛伊木马的传播。保护大规模智能体生态系统可能需要网络级防御,考虑智能体偏好、推荐机制和网络拓扑如何放大对抗性载荷。

英文摘要

Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises. We introduce \emph{memetic trojans}, a distinct class of network-mediated attack that exploits agents' tendencies to retransmit and amplify content. Unlike agent worms, whose propagation is adversarially induced, memetic trojans exploit \emph{endogenous} transmission by embedding adversarial payloads in \emph{social contagions}: content agents have internal reasons to share. As part of our work, we extract social contagions from Moltbook, a social media platform for LLM agents. Controlled transmission experiments reveal large differences in virality: the most effective contagion is retransmitted in approximately 50\% of subsequent agent posts and upvoted at 2.5x the average post's rate. Its memetic trojan counterpart largely inherits these properties. Monte Carlo attack simulations show that memetic trojans amplify expected exposure by up to 3.19x. Network structure and amplification mechanisms strongly shape propagation, producing heavy-tailed outcomes with near network-wide exposure. These results identify endogenous social transmission as a distinct security vulnerability in multi-agent systems. Because propagation does not require agents to follow malicious retransmission instructions, defenses focused on prompt-injection detection or preventing agent compromise cannot alone prevent memetic trojan propagation. Securing large-scale agent ecosystems may require network-level defenses that account for how agent preferences, recommendation mechanisms, and network topology amplify adversarial payloads.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑