发表机构
Sony Group Corporation; SiriusXM; Deezer Research; Amazon; Politecnico di Bari; Maastricht University(索尼集团公司; 天狼星卫星广播公司; Deezer研究院; 亚马逊; 巴里理工大学; 马斯特里赫特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文概述RecSys 2026挑战赛的对话式音乐推荐任务,分析16个系统在检索-重排-生成框架下的表现,指出稳健设计需结合多轮上下文、意图检测及冷启动信号,并揭示基准评估的局限性。
AI 中文摘要
RecSys挑战赛2026将对话式音乐推荐作为联合物品推荐和响应生成问题进行研究:给定多轮对话,系统必须从大型目录中检索相关曲目并生成有依据的自然语言响应。本文介绍了挑战任务、数据集、评估协议和官方结果。除了排行榜之外,我们通过通用的检索-重排-生成框架分析了16个被接受的系统,并考察了推荐性能在不同用户、请求和对话上下文中的变化。强大的系统通常结合异构候选来源,并保留来源特定的证据以进行学习重排。在系统论文和我们的组织方分析中,稳健的设计还意味着:1)将冷启动检索基于多轮对话和物品信号,2)使用意图检测器,3)建模完整的多轮上下文而不仅仅是当前查询。我们进一步指出了基准和评估协议的局限性,包括单一真实相关性和对合成对话的教师强制评估。总之,这些发现为未来的对话式推荐系统和共享评估工作提供了实用指导。
英文摘要
The RecSys Challenge 2026 studies conversational music recommendation as a joint item recommendation and response generation problem: given a multi-turn dialogue, systems must retrieve relevant tracks from a large catalog and produce a grounded natural-language response. This paper presents the challenge task, dataset, evaluation protocol, and official results. Beyond the leaderboard, we analyze the 16 accepted systems through a common retrieve--rerank--generate framework and examine how recommendation performance varies across users, requests, and dialogue contexts. Strong systems commonly combine heterogeneous candidate sources and preserve source-specific evidence for learned reranking. Across the system papers and our organizer-side analysis, robust design also means 1) grounding cold-start retrieval in multi-turn conversation and item signals, 2) using intent detectors, and 3) modeling the full multi-turn context rather than the current query alone. We further identify limitations of the benchmark and evaluation protocol, including single-ground-truth relevance and teacher-forced evaluation of synthetic dialogues. Together, these findings provide practical guidance for future conversational recommender systems and shared evaluation efforts.