发表机构
University of Zurich; Dublin City University(苏黎世大学; 都柏林城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究LLMs在推荐系统解释评估中的实用性,通过生成18种解释原型、由14种LLMs评估并与人类评分对比,发现LLMs与人类评分有中等相关性但绝对一致性低,进而给出四条实用建议。
AI 中文摘要
解释在构建可信赖的推荐系统(RS)中发挥着关键作用,但选择合适的解释方法存在挑战:现有解释生成方法繁多,却缺乏针对不同场景的最优方法指导;多数方法生成的抽象输出需进一步格式化才能便于用户使用,选项看似无穷无尽;对所有可能选项开展基于用户的评估通常不可行,而自动评估指标往往仅评估解释器的抽象输出,或需与通常不可用的基准进行比较。近期研究表明,大语言模型(LLMs)可作为解释评估的“评判者”,但其可靠性尚未得到充分探索。本文研究LLMs在为特定应用选择有效解释方法中的实用性:首先探究其在给定RS和用户不同信息时生成解释原型的能力,具体生成18种不同的解释原型,随后由14种不同规模、两种温度设置的LLMs对这些原型进行评估,并将结果与用户研究得出的人类评分进行比较。结果显示,尽管LLMs表现出与人类相似的评分模式,且与人类评分者的排名相关性达到中等水平,但其绝对评分一致性较低,且随模型规模和评估构建的不同存在显著差异。本文提出四条实用建议:保持解释生成提示简洁;选择更大规模的模型用于评估;预先测试评估构建;以及审核解释的事实准确性,因为人类和LLMs均无法可靠检测非事实内容。
英文摘要
Explanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method presents challenges. Many explanation methods exist, but little guidance exists on which is best for which setting. Existing explanation generation methods often produce abstract outputs that require further formatting to become user-friendly, with a seemingly endless pool of options. Running user-based evaluations of all possible options is usually unfeasible, while automated evaluation metrics often either assess only the explainer's abstract output or require comparison with a ground truth, which is generally unavailable. Recent studies have shown that large language models (LLMs) can serve as ``judges'' for explanation evaluation, but their reliability has not yet been thoroughly explored. This paper studies the utility of LLMs in selecting an effective explanation method for a given application. We first explore their ability to generate explanation prototypes given varying information about the RS and the user. Specifically, we generate 18 distinct explanation prototypes, which are subsequently evaluated by 14 LLMs of varying sizes across two temperature settings. We compare these against human ratings derived from a user study. Our results show that while LLMs exhibit human-like rating patterns and achieve moderate rank correlation with human raters, their absolute rating agreement is low and varies substantially by model size and evaluation construct. We derive four practical recommendations: keep explanation-generation prompts concise, prefer larger models for evaluation, pre-test evaluation constructs, and audit explanations for factual accuracy, as neither humans nor LLMs reliably detect non-factual content.
Comments10 pages, 5 figures, Accepted at ACM RecSys 2026