arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38270cs.CRcs.LGcs.MA

VirusCascade:劫持LLM驱动的推荐智能体中的协作反思

VirusCascade: Hijacking Collaborative Reflection in LLM-Powered Recommender Agents

  • Nanyang Technological University(南洋理工大学)
  • College of Computing and Data Science(计算与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

Yurong Hao, Wen Zhou, Guowei Guan, Tiantong Wu, Fuyao Zhang, Wei Yang Bryan Lim

AI总结:

针对LLM驱动推荐智能体的协作反思机制,提出首个黑盒定向推广攻击VirusCascade,利用反思持久性与跨智能体传播实现系统级劫持,在四个数据集上达到最优定向曝光效果。

AI中文摘要:

超越传统的静态评分模型,LLM驱动的智能体推荐系统(LLM-ARS)将用户和物品实例化为自主智能体,其语义状态通过称为协作反思的循环过程动态细化。虽然这一机制提高了推荐质量,但它同时引入了一个系统性漏洞:注入到单个智能体中的对抗性证据可以被合理化(rationalised)为合法的偏好叙述,写回记忆,并通过交互上下文传播给其他智能体。我们将局部的合理化过程称为反思洗白(reflection laundering),并将其通过协作反思的系统性升级称为协作反思劫持(collaborative-reflection hijacking)。现有的针对推荐系统的攻击,无论是基于交互层面的数据投毒还是文本层面的对抗扰动,都假设静态流水线,因此无法利用这种循环的、多智能体的放大路径。为弥合这一差距,我们首先进行了一项受控的漏洞分析,确立了支撑协作反思劫持的两个可利用属性:反思持久性(reflective persistence)和跨智能体传播(cross-agent propagation)。然后基于这些发现,我们提出了VirusCascade,这是首个黑盒定向推广攻击,它联合塑造语义和结构攻击面:前者确保目标物品被自然地合理化(rationalised)为满足广泛的用户偏好,后者将其定位为系统范围的传播。在四个真实世界数据集上、跨多种LLM-ARS架构的大量实验表明,VirusCascade在评估的隐蔽性约束下始终达到最先进的定向曝光效果,平均E@20达到0.384,并以+0.185的绝对差距超过最强基线。

英文摘要:

Advancing beyond traditional static scoring models, LLM-powered agentic recommender systems (LLM-ARS) instantiate users and items as autonomous agents, whose semantic states are dynamically refined through a recurrent process known as collaborative reflection. While this mechanism improves recommendation quality, it simultaneously introduces a systemic vulnerability: adversarial evidence injected into a single agent can be rationalised into a legitimate preference narrative, written back into memory, and propagated to other agents through interaction contexts. We term the local rationalisation process reflection laundering, and its system-wide escalation through collaborative reflection collaborative-reflection hijacking. Existing attacks on recommender systems, whether based on interaction-level data poisoning or text-level adversarial perturbations, assume static pipelines and thus cannot exploit this recurrent, multi-agent amplification pathway. To bridge this gap, we first conduct a controlled vulnerability analysis that establishes two exploitable properties underlying collaborative-reflection hijacking: reflective persistence and cross-agent propagation. Then building on these findings, we propose VirusCascade, the first black-box targeted promotion attack that jointly shapes semantic and structural attack surfaces: the former ensures the target item is naturally rationalised as satisfying broad user preferences, the latter positions it for system-wide propagation. Extensive experiments on four real-world datasets across diverse LLM-ARS architectures demonstrate that VirusCascade consistently achieves state-of-the-art targeted exposure under evaluated stealth constraints, reaching a mean E@20 of 0.384 and surpassing the strongest baseline by an absolute margin of +0.185.

补充信息

↑