发表机构
National Institute of Technology Kurukshetra; Queen Mary University of London(库鲁克谢特拉国家技术学院; 伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对ACM RecSys 2026挑战赛提出三阶段多模态对话音乐推荐系统,经优化后在Blind B获0.3213综合分,发现约束大语言模型注入的会话数量对Blind A的nDCG有显著影响。
AI 中文摘要
我们提出了Semiintelligence团队针对ACM RecSys 2026 TalkPlayData挑战赛的解决方案,通过多模态个性化对话推荐系统解决对话式音乐推荐问题。提交的系统采用三阶段流水线:(1)多模态检索,在七个稠密嵌入空间构建衰减加权质心——曲目和用户级CF-BPR、Qwen3(元数据、歌词、属性)、CLAP音频、SigLIP视觉,辅以BM25词汇检索和艺术家子串匹配信号,通过加权 reciprocal rank fusion(RRF)融合,权重经优化;(2)轻量级重排序(历史过滤、流行度平滑、目录多样性惩罚);(3)使用GPT-4o-mini生成角色多样化响应。除提交配置外,我们报告了开发阶段的额外组件实验——约束大语言模型引导的艺术家注入、专辑延续信号、XGBoost LambdaMART、更优的GPT-4.1响应提示,因成本和复杂度限制未部署到Blind B。我们通过差分进化在500个会话的开发集上优化RRF权重,使MRR提升19.5%。在Blind A上,54个会话的无约束大语言模型引导注入导致nDCG灾难性下降18.9%,仅9个会话的保守注入获得Blind A观测到的最佳nDCG——该发现作为Blind A观测结果需进一步验证。提交系统在Blind B上的综合得分为0.3213。
英文摘要
We present Team Semiintelligencn's solution for the ACM RecSys 2026 TalkPlayData Challenge, addressing conversational music recommendation through a multi-modal and personalized conversational recommender system. Our submitted system employs a three-stage pipeline: (1) multi-modal retrieval constructing decay-weighted centroids across seven dense embedding spaces - track- and user-level CF-BPR, Qwen3 (metadata, lyrics, attributes), CLAP audio, and SigLIP visual - supplemented by BM25 lexical retrieval and an artist substring-match signal, all fused via weighted Reciprocal Rank Fusion (RRF) with optimized signal weights; (2) lightweight reranking (history filtering, popularity smoothing, and catalog diversity penalization); and (3) persona-diversified response generation using GPT-4o-mini. Beyond this submitted configuration, we report development-time experiments with additional components - constrained LLM-guided artist injection, album continuation signals, XGBoost LambdaMART, and a superior GPT-4.1 response prompt - that were not deployed to Blind B due to cost and complexity constraints. We optimize RRF weights on a 500-session development split via differential evolution, improving MRR by +19.5%. On Blind A, we observe that unconstrained LLM-guided injection across 54 sessions causes catastrophic nDCG regression (-18.9%), while conservative injection on only 9 sessions yields the best observed Blind A nDCG - a finding we present as a Blind A observation warranting further validation. The submitted system achieves a Blind B composite score of 0.3213.