arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义与行为信号的鲁棒融合用于个性化搜索中的LLM重排序

Robust Fusion of Semantic and Behavioural Signals for LLM Reranking in Personalised Search

Aleksandr V. Petrov, Nathan Stein, Erik Lybecker, Emma Schüldt, Daniel Lazarovski, Hugues Bouchard, Mounia Lalmas

arXiv 2609.25825首次发表:更新:

发表机构

Spotify(Spotify)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对个性化搜索中LLM重排序对行为信号过度依赖的问题,提出确定性双样本特征丢弃训练,在保留QSS增益的同时提升鲁棒性,离线提升13.3%,在线提升约2%。

AI 中文摘要

个性化搜索必须在满足查询意图的同时融入用户上下文和历史交互。基于LLM的交叉编码器提供了单一的重排序接口,但将预测性行为统计注入其提示中可能鼓励捷径学习:依赖历史信号而牺牲能够泛化到稀疏或未见搜索的语义和用户上下文模式。我们在一个大型音频流媒体平台的个性化搜索系统中研究了这一问题,使用查询切片统计(QSS),一种从交互中提取的行为特征,总结了查询-候选对的历史成功情况。当该特征可用时,朴素的QSS注入提高了排序质量,但当其被移除时则降低了鲁棒性。我们通过确定性双样本特征丢弃训练来解决这一问题,该训练将每个样本呈现两次,一次包含QSS,一次不包含QSS。离线时,QSS注入在可用时将排序质量提高了13.3%。双样本训练保留了这些增益,同时在QSS移除评估下相对于朴素QSS训练将性能提高了4.0%。在在线实时测试中,两种QSS感知变体均将搜索成功率提高了约2%。聚合测试无法区分双样本训练与仅特征训练;冷启动比较与离线结果方向一致。因此,成对的包含特征和不包含特征的训练可以减少在利用强行为统计与在不可用时保持鲁棒性之间的张力。

英文摘要

Personalised search must satisfy query intent while incorporating user context and historical interactions. LLM-based cross-encoders provide a single reranking interface, but injecting predictive behavioural statistics into their prompts can encourage shortcut learning: reliance on historical signals at the expense of semantic and user-context patterns that generalise to sparse or unseen searches. We study this problem in the personalised search system of a large-scale audio streaming platform using Query Slice Stats (QSS), an interaction-derived behavioural feature summarising historical success for query-candidate pairs. Naive QSS injection improves ranking when the feature is available but reduces robustness when it is removed. We address this with deterministic dual-sample feature-dropout training, which presents each example once with QSS included and once with QSS removed. Offline, QSS injection improves ranking quality by 13.3% when available. Dual-sample training preserves these gains while improving performance under QSS-removed evaluation by 4.0% relative to naive QSS training. In a live online test, both QSS-aware variants improve search success by roughly 2%. The aggregate test does not distinguish dual-sample from features-only training; the cold-start comparison is directionally consistent with the offline results. Paired feature-present and feature-removed training can therefore reduce the tension between exploiting strong behavioural statistics and remaining robust when they are unavailable.

CommentsAccepted at the USRW Workshop at RecSys 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑