arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00086cs.CLcs.AI

检索、评分与解码影响基于大语言模型的对话式推荐系统的性能与稳定性

Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation

  • Infobip(英福比普(Infobip))

机构由 AI 辅助整理,请以论文原文为准。

Ante Kapetanovic, Tomislav Duricic, Andro Mercep, Emanuel Lacic

AI总结:

该研究在ReDial基准上对比LLM重排序器与传统基线,发现检索、评分、解码等配置会显著影响对话式推荐系统的性能与稳定性,强调需将这些配置作为必填报告项。

AI中文摘要:

大语言模型(LLM)正越来越多地被用作对话式推荐系统中的重排序器,但其测得的性能提升高度依赖于检索与推理协议。在ReDial对话式电影推荐基准上,我们在共享的“先检索后重排序”流程中,将专有模型、开放权重模型及微调后的LLM重排序器与协同过滤及序列基线进行对比。我们调整候选池规模、第一阶段检索器及解码温度:在共享语义Top-250候选池与严格的候选感知评分下,最优专有重排序器的NDCG@10达0.1497,而最强非LLM基线为0.0939;同一重排序器在零样本生成任务中达0.2925,表明无约束评分的表面优势远大于匹配池评估。此协议下,无评估的开放权重LLM优于调优后的浅层自编码器基线。对最强专有与开放权重重排序器,从语义候选切换为协同过滤候选使NDCG@10提升超50%,说明重排序器性能对候选生成高度敏感。最优专有重排序器在温度从0升至1.0时,Top-10 Jaccard距离从0.0900升至0.1240,而平均NDCG@10变化可忽略,较弱LLM则表现出更大性能退化。这些ReDial结果支持将候选生成、候选池规模、评分策略及解码配置视为必填报告字段而非实现细节。

英文摘要:

Large language models (LLMs) are increasingly used as rerankers in conversational recommender systems, yet measured gains depend strongly on the retrieval and inference protocol. On the ReDial conversational movie recommendation benchmark, we compare proprietary, open-weight, and fine-tuned LLM rerankers with collaborative-filtering and sequential baselines in a shared retrieve-then-rerank pipeline. We vary candidate-pool size, first-stage retriever, and decoding temperature. With a shared semantic top-250 candidate pool and strict candidate-aware scoring, the best proprietary reranker reaches NDCG@10 of 0.1497, compared with 0.0939 for the strongest non-LLM baseline. The same reranker reaches 0.2925 in zero-shot generation, showing that unconstrained scoring can yield a much larger apparent advantage than matched-pool evaluation. No evaluated open-weight LLM outperforms the tuned shallow autoencoder baseline under this protocol. For the strongest proprietary and open-weight rerankers, switching from semantic to collaborative-filtering candidates raises NDCG@10 by more than 50%, showing that measured reranker performance is highly sensitive to candidate generation. For the best proprietary reranker, raising temperature from 0 to 1.0 increases top-10 Jaccard distance from 0.0900 to 0.1240 while mean NDCG@10 changes negligibly, whereas weaker LLMs show larger degradation. These ReDial results support treating candidate generation, candidate-pool size, scoring policy, and decoding configuration as required reporting fields rather than implementation details.

补充信息

↑