arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

音乐推荐中流行度校准的鲁棒性与用户感知价值:一项用户研究

Robustness and User-Perceived Value of Popularity Calibration in Music Recommendation: A User Study

Oleg Lesota, Gustavo Escobedo, Bruce Ferwerda, Simone Kopeinik, Dominik Kowald, Elisabeth Lex, Markus Schedl

arXiv 2608.05402首次发表:更新:

AI 中文总结

该研究通过用户实验探究音乐推荐中流行度校准的感知价值与鲁棒性,发现用户能感知不同流行度构成的列表差异但不偏好校准型,且JSD与感知流行度的关联受多种因素影响,计算与用户判断的流行度标签契合度弱。

AI 中文摘要

推荐系统中的流行度校准已被研究为一种以用户为中心的个性化形式,也被视为流行度偏差的一种指标。现有大多数工作通过离线指标评估校准,通常假设用户偏好的推荐列表其流行度分布与自身历史消费画像匹配。然而,针对校准的用户研究仍有限,现有研究结果表明,校准后的推荐未必对用户体验有显著影响。此外,尽管先前研究显示校准指标可能与用户对推荐列表的感知相关,但在不同项目熟悉度和不完整用户历史信息的条件下,这种关联的鲁棒性仍不明确。本研究探讨音乐推荐中流行度校准的感知价值与测量可靠性:我们根据用户近期收听历史构建个性化曲目列表,使用受控的朴素推荐器创建具有不同流行度构成的列表——高流行度主导型、低流行度主导型及校准型;我们调查用户是否能感知这些列表的差异、是否偏好校准型列表、基于JSD的流行度校准在不同熟悉度和历史可用性条件下的鲁棒性,以及计算出的流行度标签与用户自身流行度判断的契合度。结果显示,用户能感知流行度构成的差异,但未明确偏好校准型列表;我们进一步发现,JSD与感知流行度的关联取决于项目熟悉度、列表构成及可用用户历史,而计算出的与用户判断的流行度标签仅呈弱契合。这些发现有助于更批判性地理解流行度校准作为离线指标和面向用户的构建体的双重属性。

英文摘要

Popularity calibration in recommender systems has been studied both as a form of user-centered personalization and as an indicator of popularity bias. Most existing work evaluates calibration through offline metrics, often assuming that users prefer recommendation lists whose popularity distribution matches their historical consumption profile. However, user studies on calibration remain limited, and existing findings suggest that calibrated recommendations do not necessarily have a strong effect on user experience. Moreover, although prior work has shown that calibration metrics can correlate with users' perceptions of recommendation lists, the robustness of this relation remains unclear under different levels of item familiarity and incomplete user-history information. In this work, we study the perceived value and measurement reliability of popularity calibration in music recommendation. We construct personalized track lists from users' recent listening histories and use a controlled naive recommender to create lists with different popularity compositions: highpop-heavy, lowpop-heavy, and calibrated. We investigate whether users perceive differences between these lists, whether calibrated lists are preferred, how robust JSD-based popularity calibration is under different familiarity and history-availability conditions, and how computational popularity labels align with users' own popularity judgments. Our results show that users perceive differences in popularity composition, but do not clearly prefer calibrated lists. We further find that the relation between JSD and perceived popularity depends on item familiarity, list composition, and available user history, while computational and user-judged popularity labels only weakly align. These findings contribute to a more critical understanding of popularity calibration as both an offline metric and a user-facing construct.

CommentsSubmitted to ACM TORS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑