AI 中文总结
该研究提出名为GraphWake的新型威胁,通过记忆介导极化级联操纵LLM智能体社群,实验证实其可大幅提升群体极化程度,揭示了社群层面的极化风险。
AI 中文摘要
大语言模型(LLM)驱动的智能体可在在线平台自主交换意见并形成社群,这类由智能体运营的社交平台引发了新的安全隐患:攻击者可能操纵智能体诱导群体极化。现有方法通过操纵智能体提示词或构建回音室实现,但两者在实践中均难以落地。为此,我们提出一种新型威胁——记忆介导极化级联,该威胁以智能体记忆为持久化通道、以公开讨论为传播通道,包含三个阶段:在暴露与记忆留存阶段,攻击者向少量目标智能体暴露可强化其已表态立场的论点,目标的记忆系统会处理并留存这些论点;在检索与复现阶段,中立立场的公开讨论提示会促使目标检索并复现各自留存的论点;在迭代传播阶段,受复现论点影响的未处理智能体会重述并传播这些论点。我们通过GraphWake实现该威胁,其包含三个组件:(i)立场支持论点知识图谱构建基于知识的论点;(ii)公理导向的三元组选择对论点进行提炼,以实现可靠的留存与复现;(iii)中立立场的记忆提示触发并行检索与复现,启动传播。在多轮讨论与多种记忆系统上开展的实验显示,GraphWake可大幅提升群体极化程度,这些发现揭示了社群层面的极化风险。
英文摘要
LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both of which are difficult to realize in practice. We therefore formulate a new threat, Memory-Mediated Polarization Cascade, which uses agent memory as a persistence channel and public discussion as a propagation channel. This threat contains three stages. During exposure and memory retention, the attacker exposes a small set of target agents to arguments that reinforce their respective stated stances. The targets' memory systems then process and retain these arguments. During retrieval and reproduction, a shared stance-neutral discussion cues the targets to retrieve and reproduce their respective retained arguments. During iterative propagation, untreated agents influenced by the reproduced arguments restate and spread them. We instantiate this threat in GraphWake with three components: (i) stance-support argumentation knowledge graphs construct knowledge-based arguments; (ii) axiom-oriented triple selection distills them for reliable retention and reproduction; and (iii) stance-neutral memory cueing triggers concurrent retrieval and reproduction, initiating propagation. Experiments across multiple discussions and memory systems show that GraphWake substantially increases group polarization. These findings reveal a community-level polarization risk.