SIREN(诱使大语言模型陷入困境):网页增强型大语言模型推荐系统中基于配对驱动的偏好操纵
SIREN (Luring LLMs onto the Rocks): PAIR-Driven Preference Manipulation in Web-RAG Recommenders
浏览论文内容
中文总结 AI 辅助
研究网页增强型大语言模型推荐系统中排序推荐的对抗操纵问题,提出SIREN方法,利用23种内容中毒技术迭代编辑检索到的网页,在两个Claude模型上实验,使选定实体多次达排名第1,验证了声明性排名等方式的有效性。
中文摘要 AI 辅助
本文研究了对网页增强型大语言模型(LLMs)生成的排序推荐进行对抗性操纵。当大语言模型通过检索和读取实时网页来回答推荐查询时,它就充当了推荐系统,每个检索到的页面都成为潜在的攻击面。先前的工作研究了虚假产品、检索中毒和排名提升。但这些研究没有比较在周围源集不变的情况下,对已检索页面的不同编辑如何改变模型的最终排名。为填补这一空白,我们提出了SIREN,这是一种自动化的攻击者 - 判断方法,它将PAIR越狱循环应用于竞争性排名操纵,目标是在大语言模型生成的推荐中将选定实体提升到排名第1。SIREN使用Anthropic的网络工具检索和捕获网页,然后使用23种内容中毒技术的可解释分类法迭代编辑检索到的源。自定义RAG重放平台保持相同的源顺序不变,因此模型排名的变化可以与提供内容的变化相关联,而不是与检索差异相关联。在两个生产Claude模型中,在八个查询 - 模型上下文中嵌套的124次技术试验中,SIREN在62次试验中使目标实体达到排名第1。然后在新会话中测试达到排名第1的有效载荷,平均成功率为0.805。在评估的设置中,声明性排名声明和种子列表通常比指令形式的注入更有效,尽管这种差异的强度取决于目标模型。据我们所知,这是首批在生产大语言模型中对竞争性排名操纵进行的对照研究之一,其中提供的源上下文保持固定。
英文摘要
This paper investigates the adversarial manipulation of the ranked recommendations produced by web-augmented large language models (LLMs). When an LLM answers a recommendation query by retrieving and reading live webpages, it acts as a recommender, and each retrieved page becomes a potential attack surface. Prior work has examined fabricated products, retrieval poisoning, and rank promotion. However, these studies do not compare how different edits to an already retrieved page change the model's final ranking while the surrounding source set remains unchanged. To address this gap, we propose SIREN, an automated attacker--judge method that adapts the PAIR jailbreaking loop to competitive rank manipulation, with the goal of moving a chosen entity to rank~1 in an LLM-generated recommendation. SIREN retrieves and captures webpages using Anthropic's web tools, then iteratively edits a retrieved source using an interpretable taxonomy of 23 content-poisoning techniques. The custom-RAG replay platform keeps the same sources in the same order, so changes in the model's ranking can be linked to changes in the supplied content rather than to differences in retrieval. Across two production Claude models, SIREN reaches rank~1 in 62 of 124 technique trials nested within eight query--model contexts. The payloads that reached rank~1 were then tested in fresh sessions, where they reproduced the result with a mean success rate of 0.805. Across the evaluated settings, declarative ranking claims and seeded lists were generally more effective than directive-form injections, although the strength of this difference depended on the target model. To the best of our knowledge, this is among the first controlled studies of competitive rank manipulation in production LLMs where the supplied source context is kept fixed.
发表机构
- The University of Queensland(昆士兰大学)
- Curtin University(科廷大学)
机构由 AI 辅助整理,请以论文原文为准。