为确认而提问:用于自信多轮大语言模型推荐的信息交互
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
浏览论文内容
中文总结 AI 辅助
该研究针对多轮LLM推荐中有效挖掘用户偏好的挑战,提出以推荐熵降为奖励的方法微调LLM,结合SFT与DPO在INSPIRED、ReDial数据集上提升了推荐质量与对话效率。
中文摘要 AI 辅助
近期大语言模型(LLM)的进展已使其可被用作对话式推荐系统(CRS),展现出强大的推荐准确率与自然对话能力。然而,引导多轮交互以有效挖掘用户偏好仍具挑战。现有方法要么使用带模板交互的独立强化学习智能体,要么优化由另一LLM评判的交互性,却未衡量实际获取的有用信息量。我们提出一种新方法,通过推荐结果的熵来衡量助手不确定性的降低量,以此量化每次交互的有效性。我们将这种熵降低作为奖励——不依赖现实场景中常不可得的真实推荐结果——来微调LLM,使其能生成策略性交互。在INSPIRED和ReDial数据集上,结合监督微调(SFT)与直接偏好优化(DPO)的实证结果显示,我们的方法同时提升了推荐质量与对话效率。
英文摘要
Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue. However, guiding multi-turn interactions to elicit user preferences effectively remains challenging. Existing approaches either use separate reinforcement learning agents with templated interactions or optimize for interactivity judged by another LLM, without measuring how much useful information is actually gained. We propose a new approach that quantifies the effectiveness of each interaction by the reduction in the assistant's uncertainty, measured via entropy over recommendations. We apply this entropy reduction as a reward---without relying on ground-truth recommendations, which are often unavailable in real-world scenarios---to fine-tune the LLM, enabling strategic interaction generation. Empirical results with supervised fine-tuning (SFT) and direct preference optimization (DPO) on the INSPIRED and ReDial datasets show that our method improves both recommendation quality and conversational efficiency.
发表机构
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。