一对一排名器中的赔率偏移滑移:诊断与修复重加权引起的Top-K错误
Odds-Shift Slippage in One-vs-Rest Rankers: Diagnosing and Repairing Reweighting-Induced Top-K Errors
浏览论文内容
中文总结 AI 辅助
本研究诊断一对一排名器中重加权导致的赔率偏移滑移,量化其影响,并提出逐标签保序回归修复方法,显著恢复Top-K性能。
中文摘要 AI 辅助
一对一排名器向每个用户展示许多稀有标签中的前$K$个,通常通过每个标签的正类权重scale_pos_weight $= n_-/n_+$来应对不平衡。Elkan恒等式表明,这样的权重将标签$j$的对数赔率偏移$\n w_j$,因此模型按加权赔率而非对precision@$K$贝叶斯最优的边际概率进行排序,并建议事后反转偏移;有限学习器如何处理数千量级的权重,以及哪种修复有效,尚未被测量。我们将承诺偏移与实际实现偏移之间的差距称为赔率偏移滑移,并在仅权重不同的LightGBM和MLP模型匹配对上测量它。在Santander数据集上,权重将MAP@7从0.808降至0.117;对于提升树对,理想赔率偏移解释了该损失的23%(在Instacart上为32%;在相同行上的MLP对为98%),其余为滑移。我们证明,叶子步长上限为$c$的提升器在$T$轮中以速率$\n$最多实现$T\n c$纳特的偏移,这通过上限扫描得到确认,并表明没有上限时,饱和单元恰好平局于1.0,超出任何可分离映射的能力。因此,解析反转仅在偏移已实现且未饱和处有效,而逐标签保序回归将Santander模型恢复到0.784(在验证期间拟合校准器时为0.780),但仅当没有校准正样本的标签映射到其先验而非直接通过时。在11个公开MULAN基准和5个学习器上,加权模型在55个单元中的8个损失超过其MAP@$K$的一半,在delicious和Corel5k上,相同的修复将其恢复到未加权水平;逐标签校准在正样本稀少时有害,交叉验证规则可消除此危害。该方案以oddslip发布。
英文摘要
One-vs-rest rankers that show each user the top-$K$ of many rare labels usually counter imbalance with a per-label positive-class weight, scale_pos_weight $= n_-/n_+$. Elkan's identity says such a weight shifts label $j$'s log-odds by $\ln w_j$, so the model ranks by weighted odds rather than by the marginal that is Bayes-optimal for precision@$K$, and suggests inverting the shift afterwards; what a finite learner does with a weight in the thousands, and which repair then works, has not been measured. We call the gap between the promised and the realized shift odds-shift slippage and measure it on matched pairs of LightGBM and MLP models that differ only in the weights. On Santander the weight takes MAP@7 from 0.808 to 0.117; for the boosted pairs the ideal odds shift accounts for 23% of that loss (32% on Instacart; 98% for an MLP pair on the same rows) and slippage for the rest. We prove that a booster whose leaf steps are capped at $c$ realizes at most $Tηc$ nat of shift in $T$ rounds at rate $η$, which a cap sweep confirms, and show that without a cap saturated cells tie at exactly 1.0, beyond the reach of any separable map. The analytic inversion therefore pays only where the shift was realized and nothing saturated, whereas per-label isotonic regression returns the Santander model to 0.784 (0.780 with the calibrator fitted on the validation period), but only if labels without calibration positives are mapped to their prior rather than passed through. On 11 public MULAN benchmarks and 5 learners the weighted model loses more than half of its MAP@$K$ in 8 of 55 cells, and on delicious and Corel5k the same repair returns it to the unweighted level; per-label calibration hurts where positives are scarce, a harm that a cross-validated rule removes. The recipe is released as oddslip.
发表机构
- Shiga University(滋贺大学)
机构由 AI 辅助整理,请以论文原文为准。