发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型列表式推荐器中因位置偏差产生的顺序敏感性攻击面,引入“promo@k”量化漏洞,通过实验得出不同方法对其的影响,证明输入候选顺序是安全相关攻击向量。
AI 中文摘要
在推荐系统中用作列表式重排器的大语言模型在将候选集序列化到提示中时存在位置偏差。我们表明这种顺序敏感性会产生一个可利用的攻击面:攻击者可以仅通过重新排列候选项,而不改变项目内容、标签或模型参数,就将标签为0的目标提升到前k名。我们引入“promo@k”来量化此漏洞,衡量通过排列可提升到前k名的标签为0的目标的比例。在三个领域(MovieLens、亚马逊图书和亚马逊时尚)进行评估,在攻击预算R = 50次排序时,“promo@5”高达0.57。此外,普通排列稳定性无需运行攻击即可预测漏洞。虽然双向T5编码器评分器可降低暴露风险,但排列一致性正则化和架构不变性可有效缓解此问题。逐点评分可避免偏差问题,但会降低排序质量。这些结果表明,列表式大语言模型重排中的输入候选顺序是一个与安全相关的攻击向量。代码和数据可在该https网址获取。
英文摘要
Large language models (LLMs) used as listwise rerankers in recommendation systems suffer from position bias when serializing candidate sets into prompts. We show this order sensitivity creates an exploitable attack surface: an attacker can promote a label-0 target into the top-$k$ solely by reordering candidates, without changing item content, labels, or model parameters. We introduce $\mathrm{promo}@k$ to quantify this vulnerability, measuring the fraction of label-0 targets that can be elevated into top-$k$ rankings via permutation. Evaluating across three domains (MovieLens, Amazon Books, and Amazon Fashion), $\mathrm{promo}@5$ reaches up to 0.57 at an attack budget of $R$ = 50 orderings. Furthermore, ordinary permutation stability predicts vulnerability without running the attack. While a bidirectional T5 encoder scorer reduces exposure, permutation-consistency regularization and architectural invariance effectively mitigate it. Pointwise scoring avoids the bias issue but degrades ranking quality. These results demonstrate that input candidate order in listwise LLM reranking is a security-relevant attack vector. Code and data are available at https://github.com/geoz-lab/position_bias_attack.
Comments13 pages, 6 figures, 11 tables. Code and data available at https://github.com/geoz-lab/position_bias_attack