AI 中文总结
针对推荐评估中平局打破规则影响top-k结果的问题,通过顺序不变性审计发现不同平局策略导致NDCG大幅波动,并提出报告清单以规范评估。
AI 中文摘要
离线top-k评估通常将一个留出的相关项与采样负样本一起排序。当多个候选获得完全相同的分数时,平局打破规则成为排序的一部分。一种常见实现是将相关项存储在前,然后应用稳定排序,这会在相同分数之间保留输入顺序;因此相关项在每次平局中获胜。当对输入候选进行排列而不改变其身份、标签或分数时,最终排序保持不变,我们称这样的评估器为行顺序不变的。我们通过固定候选和分数,仅改变平局打破规则来审计这一属性。在30,000条Amazon Beauty & Personal Care数据上,对于评分加权的属性重叠分数,在输入顺序平局打破下NDCG@10为0.85。基于用户和物品ID的确定性哈希平局打破将其降至0.17。在均匀随机平局打破下的精确期望与100个独立哈希种子的平均值紧密匹配,而具有少量精确平局的残差属性分数几乎不变。MovieLens Tag Genome对属性重叠分数显示出相同模式,而物品流行度几乎不变。我们推导了当相关项在具有相同分数的候选中随机排序时,在截止k处的期望命中率和NDCG,并提供了实用的报告清单。只要精确平局影响top-k成员资格或排名,同样的问题可能出现在采样或全目录评估中。
英文摘要
Offline top-k evaluation often ranks one held-out relevant item together with sampled negatives. When several candidates receive exactly the same score, the tie-breaking rule becomes part of the ranking. A common implementation stores the relevant item first and then applies a stable sort, which preserves input order among equal scores; the relevant item therefore wins every tie. We call an evaluator row-order invariant when permuting the input candidates without changing their identities, labels, or scores leaves the final ranking unchanged. We audit this property by holding candidates and scores fixed and changing only the tie-breaking rule. On 30,000 Amazon Beauty & Personal Care rows, NDCG@10 for a rating-weighted attribute-overlap score is 0.85 under input-order tie-breaking. A deterministic hash tie-break based on user and item IDs lowers it to 0.17. The exact expectation under uniform random tie-breaking closely matches the mean over 100 independent hash seeds, while a residualized attribute score with few exact ties is nearly unchanged. MovieLens Tag Genome shows the same pattern for an attribute-overlap score, whereas item popularity is nearly unchanged. We derive expected Hit Rate and NDCG at cutoff k when the relevant item is randomly ordered among candidates with the same score, and we provide a practical reporting checklist. The same issue can occur in sampled or full-catalog evaluation whenever exact ties affect top-k membership or rank.
Comments8 pages, 3 tables. Accepted at FRAME'26: Methodology First - Rethinking Research Assessment in RecSys Workshop, co-located with ACM RecSys 2026