AI 中文总结
研究对话式音乐推荐系统中用大语言模型评估系统响应的可靠性,通过抽取会话、用四个指令微调的大语言模型生成候选响应,收集领域专家评分并分析,发现其与人工评估适度正相关且优于基线,还分析了性能随模型规模等的变化。
AI 中文摘要
对话式推荐系统旨在实现推荐相关项目和生成自然语言响应这两个主要目标。推荐准确性可通过既定排名指标有效衡量,而响应生成的评估则面临更根本的挑战。尽管人工评估仍是金标准,但成本和可扩展性限制促使采用大语言模型作为评判器这一有前景的替代方法,其在对话式推荐系统中与人工判断的一致性仍是未决问题。本文进行了首次用户研究,以实证评估大语言模型作为评判器评估对话式推荐系统响应的可靠性。我们抽取20个多轮音乐推荐会话,使用四个指令微调的大语言模型生成候选系统响应,在模型规模上诱导响应质量的差异。我们从20位领域专家注释者那里收集了400个评分,他们从个性化质量和解释质量两个维度评估每个响应。通过自举相关分析,我们发现基于大语言模型的评判器与人工评估呈现适度正相关,且优于所有基于参考的基线。此外,我们分析了评判器性能如何根据模型规模和条件信息而变化,为部署大语言模型作为评判器提供了实用指导。
英文摘要
Conversational Recommendation Systems (CRS) aim to achieve two primary objectives: recommending relevant items and generating natural language responses. While recommendation accuracy is effectively measured by established ranking metrics, the evaluation of response generation poses a more fundamental challenge. Although human evaluation remains the gold standard, its cost and scalability constraints have motivated the adoption of LLM-as-a-judge as a promising proxy, whose alignment with human judgment in the context of CRS remains an open question. In this paper, we present the first user study to empirically assess the reliability of LLM-as-a-judge for evaluating CRS responses. We sample 20 multi-turn music recommendation sessions and generate candidate system responses using four instruction-tuned LLMs, inducing variance in response quality across model scales. We collect $n{=}400$ ratings from 20 domain-expert annotators, who evaluate each response across two dimensions: Personalization Quality and Explanation Quality. Through bootstrapped correlation analysis, we find that LLM-based judges exhibit moderate positive alignment with human assessments and outperform all reference-based baselines. Furthermore, we analyze how judge performance varies according to model scale and conditioning information, providing practical guidance for deploying LLM-as-a-judge.
CommentsAccepted for publication at the 20th ACM Conference on Recommender Systems (ACM-RecSys 2026)