发表机构
Institute of Intelligent Computing, University of Electronic Science and Technology of China(电子科技大学智能计算研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PromptShift框架,无需训练即可量化并缓解LLM推荐系统中身份线索引起的偏好漂移,通过漂移度量与重排序策略,显著降低漂移并提升推荐质量。
AI 中文摘要
在大语言模型(LLM)为基础的推荐系统中,即使底层行为证据保持不变,提示中嵌入的身份线索也可能将推荐结果引向群体层面的模式。我们提出了PromptShift,一个可解释的、无需训练的框架,用于量化和缓解此类身份线索偏好漂移。我们将漂移(Drift)定义为在身份线索提示下生成的推荐列表与仅基于同一用户交互历史生成的参考列表之间,在条目成员和排序顺序两方面的分歧。SliceShift则衡量相对于仅基于历史的参考列表,线索化列表向在身份切片内比在全局用户群体中更受欢迎的条目倾斜的程度。除了传统的准确性指标,我们提出了DifHitRate,一个难度加权的命中指标,只对相关条目给予信用,对在身份切片内不太受欢迎且在列表中排名更高的命中赋予更高的信用。所有组件都由一个基于正交互构建的身份切片-条目表支持,该表还支持一种自适应的后处理重排序策略:重排序器在原始LLM排序和逆切片流行度之间进行插值,并使用个性化的插值权重。在三个LLM上的两个数据集上的实验表明,身份线索提示比无身份线索的释义控制产生更高的平均漂移,这一效应超出了通用措辞敏感性的范畴,并且SliceShift在所有六个数据集-模型设置中均为正值。PromptShift持续降低漂移和SliceShift,将宏平均SliceShift降低了62.42%,同时提高了DifHitRate、HitRate和MRR。这些结果表明,身份线索偏好漂移可以在不进行任何模型训练的情况下被测量和缓解,尽管会带来适度的、依赖指标的效用成本。
英文摘要
In large language model-based recommender systems, identity cues embedded in prompts can steer recommendations toward group-level patterns even when the underlying behavioral evidence remains unchanged. We introduce PromptShift, an interpretable, training-free framework for quantifying and mitigating such identity-cue preference drift. We define Drift as the divergence, in both item membership and ranking order, between a recommendation list generated under an identity-cued prompt and the reference list produced from the same user's interaction history alone. SliceShift then measures the extent to which a cued list gravitates, relative to the history-only reference, toward items that are more popular within the cued slice than among the global user population. Beyond conventional accuracy, we propose DifHitRate, a difficulty-weighted hit metric that credits only relevant items, assigning higher credit to hits that are less popular within the cued slice and ranked higher in the list. All components are supported by an identity-slice-by-item table constructed from positive interactions, which further enables an adaptive post-hoc reranking strategy: the reranker interpolates between the original LLM ranking and inverse slice-popularity, with personalized interpolation weight. Experiments on two datasets with three LLMs show that identity-cued prompts incur higher mean Drift than identity-free paraphrase controls, an effect beyond generic wording sensitivity, and that SliceShift is positive across all six dataset-model settings. PromptShift consistently reduces both Drift and SliceShift, lowering macro-mean SliceShift by 62.42%, while improving DifHitRate, HitRate and MRR. These results demonstrate that identity-cue preference drift can be measured and mitigated without any model training, albeit with a modest, metric-dependent utility cost.
Comments11 pages, 1 figure