语义ID推荐器的离线策略评估:模型自身的代码层次结构是否有帮助?
Off-Policy Evaluation for Semantic ID Recommenders: Does the Model's Own Code Hierarchy Help?
查看机构详情
- Criteo AI Lab(Criteo人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究探讨能否用生成式推荐器的语义ID树作为离线策略评估的动作抽象,提出通过代码前缀簇粗化的方法提升评估效果,明确粗化增益的来源及分辨率深度的调节作用。
中文摘要 AI 辅助
生成式推荐器越来越多地输出语义ID(Semantic IDs, SIDs):每个物品是来自残差量化器的分层离散代码的短序列,通过自回归方式解码。在投入宝贵的A/B测试资源之前,团队可能需要离线评估哪些解码器或重排序变体值得测试——这正是离线策略评估(Off-Policy Evaluation, OPE)的任务。我们提出一个简单问题:模型自身的SID树能否作为该OPE的动作抽象?我们的回答分为三部分:(i)在实际推荐器使用的近argmax日志策略下,基于物品的OPE是不可行的,因为生产日志中物品级有效样本量通常很小,但将物品边缘化到代码前缀簇可恢复可估计的支持并减少误差;(ii)这种增益源于 coarsening(粗化),而非特定的层次结构,但SID树使得在生成式系统中粗化可行——每个簇的质量可由解码器精确且廉价地返回,而平面聚类需要枚举仅基于代码的解码器未直接暴露的物品/叶子质量;(iii)分辨率深度是关键调节参数,在支持稀缺时采用更粗的粒度,且条件偏差界将粗化偏差与量化器的最坏情况重构残差、目标-日志分布差异关联起来。
英文摘要
Generative recommenders increasingly emit semantic IDs (SIDs): each item is a short sequence of hierarchical discrete codes from a residual quantizer, decoded autoregressively. Before spending scarce A/B-test, a team may decide offline which decoder or reranking variants are worth testing - a job for off-policy evaluation (OPE). We ask a simple question: can the model's own SID tree serve as the action abstraction for that OPE? Our answer has three parts. (i) Under the near-argmax logging real recommenders use, per-item OPE is hopeless - as item-level effective sample size is usually small on production logs - but marginalizing items to code-prefix clusters restores estimable support and cuts error. (ii) This gain is thanks to coarsening, not to the hierarchy specifically; but the SID tree is what makes coarsening feasible in a generative system - each cluster's mass is exactly and cheaply returned by the decoder, whereas flat clustering requires enumerating item/leaf masses that a code-only decoder does not directly expose. (iii) Resolution depth is the operative knob - coarser under scarce support - and a conditional bias bound links the coarsening bias to the quantizer's worst-case reconstruction residual and the target-logging divergence.