arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29578cs.IR

选择导出的物品图的边谱:强边和弱边在协同过滤中编码不同的关系

The Edge Spectrum of Choice-Derived Item Graphs: Strong and Weak Edges Encode Different Relations in Collaborative Filtering

  • Hokkaido University(北海道大学)

机构由 AI 辅助整理,请以论文原文为准。

Keigo Sakurai, Takahiro Ogawa, Miki Haseyama

AI总结:

该研究发现选择导出的物品图中强边与弱边编码不同关系,将诊断转化为可复用协议,为协同过滤中选择导出算子的应用提供指导。

AI中文摘要:

图协同过滤依赖于物品-物品图,其边用于正向平滑,隐含假设是更强的边编码与较弱边相同的关系。我们表明,对于一类具有实际重要性的图,该假设不成立:那些边权重来自选择模型的图。在这类图上,强边和弱边编码性质不同的关系,我们称之为边谱。具体而言,强边集中在点击物品的 slate 内竞争对手上,恰好是 slate 内排名梯度将其推开的那些对,而弱边则不然。我们将此形式化为平滑算子与排名梯度之间的符号不匹配,并证明 co-click 图由于其构造不会表现出相同的错位。该诊断解释了在 MIND 和 EB-NeRD 上的三个经验观察:(i)即插即用的选择导出算子无法击败 co-click,尽管它们索引的邻域在结构上不同;(ii)统一标量修正(符号翻转、slate 内边际损失)可预测地失败,因为错位存在于图中,而非损失中;(iii)只有具有边缘幅度感知的算子(其 regime 边界由诊断而非调优确定)才能恢复预测的排序。因此,邻居截断 k 是语义开关,而非稀疏化超参数。我们的主张涉及哪些干预措施会失败或成功及其原因,而非绝对的 headline 收益,诊断本身预测在我们观察到的衰减传播通道下,绝对收益会很小。我们将该诊断转化为从业者可在部署任何选择导出的物品侧算子之前运行的可复用协议。代码:this https URL。

英文摘要:

Graph collaborative filtering relies on item--item graphs whose edges are used for positive smoothing, under the implicit assumption that stronger edges encode more of the same relation as weaker ones. We show that this assumption fails for a practically important class of graphs: those whose edge weights come from a choice model. On such graphs, strong and weak edges encode qualitatively different relations, which we call an edge spectrum. Specifically, strong edges concentrate on the in-slate competitors of clicked items, exactly the pairs that the within-slate ranking gradient pushes apart, while weak edges do not. We formalize this as a sign mismatch between the smoothing operator and the ranking gradient, and prove that co-click graphs cannot exhibit the same misalignment by construction. This diagnosis explains three empirical observations on MIND and EB-NeRD: (i) drop-in choice-derived operators do not beat co-click, despite indexing structurally distinct neighborhoods; (ii) uniform scalar fixes (sign flip, in-slate margin loss) fail predictably, because the misalignment lives in the graph, not in the loss; (iii) only edge-magnitude-aware operators, with the regime boundary located by the diagnosis rather than by tuning, recover the predicted ordering. The neighbor cutoff $k$ is therefore a semantic switch, not a sparsification hyperparameter. Our claim concerns which interventions fail or succeed and why, not absolute headline gains, which the diagnosis itself predicts to be small under the attenuated propagation channel we observe. We turn the diagnosis into a reusable protocol practitioners can run before deploying any choice-derived item-side operator. Code: https://github.com/kyomusso/Edge-Spectrum-in-CF.

补充信息

↑