arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16872cs.IR

展示份额预测:排序系统的离线评估任务

Impression Share Prediction: An Offline Evaluation Task for Ranking Systems

Mohsen Malmir, Houssam Nassif, Danish Nasir Shaikh, Taher Rahgooy, Murat Ali Bayir

首次发表
浏览论文内容

中文总结 AI 辅助

针对排序系统离线评估无法揭示展示份额变化的问题,提出展示份额预测任务,构建结构因果模型并开发统计学习框架,经实验验证可有效预测展示份额分布。

中文摘要 AI 辅助

离线评估是排序模型A/B测试中在线评估前的主要环节。标准离线指标衡量预测准确率,但仅作为下游效用的替代指标:模型可在提升这些指标的同时,以降低下游效用的方式在目标分桶间重新分配展示量。目前尚无离线方法能在在线评估前揭示这些展示份额的变化。我们提出“展示份额预测”作为一项离线评估任务:给定候选排序模型,预测其在目标分桶间产生的展示量分布——展示量按优化目标(如点击、视频观看)分组。该任务本质是反事实的,因为候选模型从未实际投放过流量。我们提出一个结构因果模型,用于描述模型预测和投放能力如何共同决定展示分配,并证明可从观测数据中识别反事实效应。在此基础上,我们开发了一个统计学习框架,基于历史数据,从候选模型的早期交互置信信号和当前系统状态预测展示份额。在多个排序模型系列的数据上,对于训练期间见过的模型,随机森林相比常数基线将L1误差降低了49%。对于按首次出现时间评估的未见过模型,首次小时是最接近真实在线评估的时段,也是最难的:由于能力状态仍反映先前模型,随机森林表现低于基线。而模拟近期拍卖动态2小时滚动的编码器条件架构,在该 regime 中恢复了22%的L1增益。

英文摘要

Offline evaluation is a major gateway before online evaluation of ranking models in A/B testing. Standard offline metrics measure predictive accuracy, but are only a surrogate for downstream utility: a model can improve them while redistributing impressions across objective buckets in ways that degrade downstream utility. No offline method surfaces these impression share shifts before online evaluation. We propose \emph{impression share prediction} as an offline evaluation task: given a candidate ranking model, predict the distribution of impressions it would produce across objective buckets - impressions grouped by optimization goal (e.g., click, video view). The task is inherently counterfactual, since the candidate has never served live traffic. We propose a structural causal model of how model predictions and delivery capacity jointly determine impression allocation, and show the counterfactual effect is identified from observational data. Building on this, we develop a statistical learning framework that predicts impression shares from a candidate's early-interaction confidence signals and current system state, trained on historical data. On data from multiple ranking model families, a Random Forest reduces L1 error by 49\% over a constant baseline for models seen during training. For held-out models, evaluated by time since first appearance, the first hour is the closest analog to true online evaluation and the hardest: the Random Forest falls below the baseline because the capacity state still reflects the prior model. An encoder-conditioned architecture that simulates a 2-hour rollout over recent auction dynamics recovers $+$22\% L1 in this regime.

补充信息

↑