发表机构
Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对基于相似度的推荐系统,提出并验证了一种利用点赞评分反馈回路的多阶段投毒攻击,可引导智能体信息流并规避检索阈值。
AI 中文摘要
随着大型语言模型(LLM)及基于LLM的智能体(agent)的最新进展,这些智能体正变得越来越自主,并在互联网上获得更广泛的权限来代表用户执行操作。然而,部署在社交媒体平台上的自动化智能体(例如,用于管理用户个人账户)的脆弱性仍未得到充分探索。现有的智能体投毒研究通常假设对手能够将投毒内容暴露给智能体。尽管这种攻击直接且有效,但它更容易被检测和缓解。在社交媒体平台的背景下,这留下了一个问题:推荐系统本身是否会以更隐蔽的方式将此类内容呈现给智能体。通过理论分析,我们展示了OASIS中使用的点赞评分机制可以被利用,并刻画了多阶段投毒帖子链能够引导智能体信息流的条件。基于这些见解,我们进一步开发了一种算法,用于制作逼真的投毒帖子。实验支持了我们的理论发现,并证明了所提出算法的有效性。值得注意的是,通过利用点赞评分反馈回路,即使投毒帖子的用户-帖子相似度低于检索阈值,该攻击也能导致推荐系统选择这些投毒帖子。
英文摘要
With recent advancements in large language models (LLMs) and LLM-based agents, these agents are becoming increasingly autonomous and gaining broader access to act on users' behalf on the internet. However, the vulnerability of automated agents deployed on social media platforms (e.g., for managing a user's personal account) remains underexplored. Existing studies on agent poisoning typically assume that the adversary can expose poisoned content to the agent. Although such an attack is direct and effective, it is more easily detected and mitigated. In the context of social media platforms, this leaves open whether the recommendation system itself would surface such content to the agent in a more subtle manner. Through theoretical analysis, we show that the like-score mechanism used in OASIS can be exploited, and we characterize the conditions under which a multi-stage chain of poisoned posts can steer the agent's feed. Based on these insights, we further develop an algorithm that crafts realistic poisoned posts. Experiments support our theoretical findings and demonstrate the effectiveness of the proposed algorithm. Notably, by exploiting the like-score feedback loop, the attack causes the recommendation system to select poisoned posts even when their user-post similarity falls below the retrieval threshold.