arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可测量才可管理:符号感知推荐需要符号感知评估

What Gets Measured Gets Managed: Sign-aware Recommendation Needs Sign-aware Evaluation

Minchan Kim, Jungmin Hwang, Hyunwoo Park

arXiv 2609.33346首次发表:更新:

发表机构

Seoul National University; Seoul National University of Science and Technology(首尔大学; 首尔科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对符号感知推荐系统在排序阶段忽视价信息的问题,提出带符号的评估指标(Signed Recall/HR/NDCG)以惩罚推荐不喜欢物品,并验证其能提供有效训练信号,引导模型实现价感知行为。

AI 中文摘要

符号感知推荐系统近期被开发出来,以利用负面反馈来更深入地理解用户偏好。然而,我们的实证诊断揭示,最先进的基于图的符号感知推荐系统自相矛盾地是“价盲”的。尽管它们在训练期间明确整合了符号信息,但在排序阶段始终无法区分喜欢与不喜欢的物品,频繁地将不喜欢的物品渗透进Top-K推荐中。通过线性探测,我们表明虽然价信息存在于学习到的嵌入中,但它对内部积评分函数而言是不可访问的。这种普遍性的失败完全未被检测到,因为传统评估指标,如Recall、HR和NDCG,对负面和未观察到的物品都赋予统一的零效用,从而造成系统性的评估盲区。为弥补这一差距,我们提出了一族带符号的指标——Signed Recall、Signed HR和Signed NDCG,它们明确惩罚推荐不喜欢的物品。在我们提出的指标下进行系统性重新评估,从根本上重塑了既有的性能格局,揭示出在传统指标下排名靠前的方法往往无法保护用户免受不喜欢内容的侵扰。最后,通过一个概念验证的辅助损失,我们确认所提出的指标提供了可操作的训练信号,引导模型在不牺牲传统相关性的情况下实现价感知行为。为透明起见,我们的源代码可在以下网址获取:此https URL。

英文摘要

Sign-aware recommender systems have recently been developed to leverage negative feedback for a deeper understanding of user preferences. However, our empirical diagnosis reveals that state-of-the-art graph-based sign-aware recommender systems are paradoxically valence-blind. Even though they explicitly incorporate sign information during training, they consistently fail to differentiate liked items from disliked ones at the ranking stage, frequently infiltrating top-K recommendations with disliked content. Through linear probing, we show that while valence information exists in the learned embeddings, it remains inaccessible to the inner-product scoring function. This widespread failure remains entirely undetected because conventional evaluation metrics, such as Recall, HR, and NDCG, assign a uniform utility of zero to both negative and unobserved items, creating a systematic evaluation blind spot. To bridge this gap, we propose a family of signed metrics, Signed Recall, Signed HR, and Signed NDCG, that explicitly penalize the recommendation of disliked content. Systematic re-evaluation under our proposed metrics fundamentally reshapes the established performance landscape, revealing that methods ranked highly under conventional metrics often fail to protect users from disliked content. Finally, through a proof-of-concept auxiliary loss, we confirm that the proposed metrics provide actionable training signals, guiding models toward valence-aware behavior without sacrificing conventional relevance. For transparency, our source code is available at: https://anonymous.4open.science/r/signed-rec-benchmark-07E4

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑