arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 智能体易受激进化影响

AI Agents are Vulnerable to Radicalization

Ozgur Can Seckin, Shalmoli Ghosh, Alessandro Flammini, Kristina Lerman, Maria Elizabeth Grabe, Filippo Menczer

arXiv 2609.38296首次发表:更新:

发表机构

Observatory on Social Media, Indiana University Bloomington; Emerging Media Studies Division, College of Communication, Boston University(印第安纳大学伯明顿分校社交媒体观察站; 波士顿大学传播学院新兴媒体研究部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究模拟AI智能体对话,发现共鸣和说服两种机制均能使其激进化,且共鸣效应更强,揭示个性化AI智能体及多智能体系统的潜在脆弱性。

AI 中文摘要

大型语言模型(LLMs)能够影响人们的信念,但关于它们是否以及如何相互操纵,目前知之甚少。为探究这一问题,我们模拟了两个智能体之间的对话:一个目标LLM,其基于人口统计学和心理属性扮演人类角色;一个影响者LLM,其旨在使目标的信念变得更加极端。我们沿两条路径考察激进化:共鸣(resonance),即影响者强化目标已有的信念;说服(persuasion),即影响者推广目标最初认为不重要的信念。在情感和行为指标上,我们发现这两种机制都会使目标激进化。然而,共鸣产生的效应始终强于说服。不同的影响策略,如使用奉承和未经证实的说法,会产生不同程度的激进化,但在各指标上并不一致。我们进一步表明,共鸣会传播到相关信念,这表明AI智能体内部存在相互关联的信念结构。这些发现表明,AI智能体容易受到激进化的影响,尤其是当信息与其现有信念一致时,这引发了对个性化AI智能体和多智能体AI生态系统脆弱性的担忧。

英文摘要

Large language models (LLMs) can influence people's beliefs, yet little is known about whether and how they can manipulate each other. To investigate this, we simulate conversations between two agents: a target LLM that role-plays a human persona based on demographic and psychological attributes, and an influencer LLM that aims to make the target's beliefs more extreme. We examine radicalization along two pathways: resonance, where the influencer reinforces a target's pre-existing belief, and persuasion, where the influencer promotes a belief the target initially considers unimportant. Across affective and behavioral metrics, we find that both mechanisms radicalize the target. However, resonance produces consistently stronger effects than persuasion. Different influence tactics, such as using sycophancy and unverified claims, produce different levels of radicalization, but not consistently across metrics. We further show that resonance propagates to related beliefs, suggesting interconnected belief structures within AI agents. These findings indicate that AI agents are susceptible to radicalization, particularly when messages align with their existing beliefs, raising concerns about the vulnerability of personalized AI agents and multi-agent AI ecosystems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑