arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03091cs.IR

位置偏差破坏基于大语言模型的列表式重排序中的偏好一致性

Position Bias Undermines Preference Consistency in Listwise LLM-Based Reranking

Ethan Bito, Yongli Ren, Estrid He

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示基于LLM的列表式重排序存在位置偏差,引入多维度评估框架验证其会破坏偏好一致性,且均衡位置曝光无法修复该问题。

中文摘要 AI 辅助

大语言模型(LLM)已成为推荐系统中颇具潜力的列表式重排序模型,但在候选项排列等价的情况下其可靠性仍不明确。由于推荐候选项构成无序集合,重排序模型不应依赖于对其进行序列化的任意顺序。然而,仅解码器式LLM重排序模型会使输入顺序影响模型得分、成对偏好及最终排名。本研究探讨位置偏差如何影响基于LLM的重排序模型的排名过程,未仅测量最终排名列表的变化,而是将等价候选项排列下产生的排名视为诱导偏好系统的观测结果。我们引入评估框架,用于测量成对偏好不稳定性、全局偏好不一致性及列表式输出一致性,该框架从成对、全局及输出层面表征候选项顺序敏感性。在多个LLM、数据集及列表长度上开展的实验表明,这些一致性指标密切相关,但可能与推荐有效性及边际位置曝光偏差存在差异。提升相关性或均衡各位置曝光并不一定能恢复稳定的成对偏好、全局一致的偏好结构或一致的排名输出。上述结果表明,减少边际曝光偏差不足以确立LLM重排序模型中排名函数的有效性。代码可在该URL获取。

英文摘要

Large language models (LLMs) have emerged as promising listwise rerankers for recommender systems, but their reliability under equivalent candidate permutations remains unclear. Since recommendation candidates form an unordered set, a reranker should not depend on the arbitrary order used to serialize them. However, decoder-only LLM rerankers can allow input order to affect model scores, pairwise preferences, and rankings. We study how position bias affects the ranking process induced by LLM-based rerankers. Instead of measuring only changes in final ranked lists, we treat rankings produced under equivalent candidate permutations as observations of an induced preference system. We introduce an evaluation framework measuring pairwise preference instability, global preference inconsistency, and listwise output consistency. This framework characterizes candidate-order sensitivity at the pairwise, global, and output levels. Experiments across multiple LLMs, datasets, and list lengths show that these consistency measures are closely aligned, but can diverge from recommendation effectiveness and marginal position-exposure bias. Improving relevance or flattening exposure across positions does not necessarily restore stable pairwise preferences, globally coherent preference structures, or consistent ranked outputs. These results show that reducing marginal exposure skew is insufficient to establish ranking-function validity in LLM-based reranking. Code is available at https://github.com/ejbito/InvariRank .

补充信息

↑