发表机构
Hong Kong Baptist University(香港浸会大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模推荐系统中 Swing 分数计算代价高的问题,提出 ASC 与 K-ASC 两种高效算法,结合随机化与过滤细化,提供理论保证,在八个真实数据集上实现数量级加速。
AI 中文摘要
给定一个用户-物品图 $G$,一个查询物品 $v_q$ 和一个目标物品 $v_t$,物品对 $(v_q, v_t)$ 的 Swing 分数 $sw(v_q, v_t)$ 利用用户-物品-用户的交互结构来评估它们的相似性。该度量在物品到物品(i2i)检索任务中被发现非常有效,并在工业级推荐系统中得到广泛应用。然而,现有的计算 Swing 分数的解决方案要么由于相对于物品度的二次时间复杂度而代价过高,要么依赖于产生不令人满意质量的截断启发式方法,这使得它们在具有数十亿交互的图上尤其不切实际。在本文中,我们提出了 ASC 和 $K$-ASC,两种新颖且高效的算法,用于近似和 top-$K$ Swing 查询,以解决上述局限性。具体来说,这些算法在 Swing 值的概率相对误差和加性误差方面提供了严格的理论保证。ASC 的基本思想是以一种简单但非平凡的方式结合两种随机算法 GNS 和 USS,以最小的运行时间成本自适应地处理高和低度查询物品。特别是,$K$-ASC 通过带有精心设计启发式方法的过滤-细化范式,为 top-$K$ 查询提供了实用的效率和有效性。在八个真实数据集上的大量实验表明,ASC 和 $K$-ASC 在计算时间上可以实现比竞争对手数量级的加速,同时提供相同的近似和 top-$K$ 查询结果质量,并且特别是,$K$-ASC 在包括十亿边 Yambda 和 MAG 数据集在内的大规模图上非常高效。
英文摘要
Given a user-item graph $G$, a query item $v_q$ and a target item $v_t$, the Swing score $sw(v_q, v_t)$ of the item pair $(v_q, v_t)$ leverages the user-item-user interaction structure to evaluate their similarity. This measure is found to be highly effective in item-to-item (i2i) retrieval task and finds extensive applications in industrial-scale recommender systems. However, existing solutions towards computing Swing scores are either prohibitively expensive due to their quadratic time complexity w.r.t. the item degree, or rely on truncation heuristics that yield unsatisfactory quality, rendering them impractical particularly on graphs with billions of interactions. In this paper, we present ASC and $K$-ASC, two novel and efficient algorithms for approximate and top-$K$ Swing queries, to address the aforementioned limitations. Specifically, these algorithms provide rigorous theoretical guarantees in probabilistic relative and additive errors of Swing values. The basic idea of ASC is to combine two randomized algorithms, GNS and USS, in a simple yet non-trivial way to adaptively process high- and low-degree query items with minimal runtime cost. In particular, $K$-ASC offers practical efficiency and effectiveness for top-$K$ queries through a filter-refinement paradigm with carefully-designed heuristics. Extensive experiments over eight real datasets demonstrate that ASC and $K$-ASC can achieve orders of magnitude speed-up over competitors in terms of computational time while offering the same approximate and top-$K$ query result quality, and in particular, $K$-ASC is highly efficient on massive graphs including the billion-edge Yambda and MAG datasets.
Comments23 pages. The technical report for the paper titled "Efficient Swing Computation for Retrieval in Large-Scale Recommender Systems" in SIGMOD 2027