arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28383cs.IR

超越标量:用于观看时长预测的分布式服务接口

Beyond a Scalar: Distributional Serving Interfaces for Watch-Time Prediction

  • Shanghai Jiao Tong University(上海交通大学)
  • Rice University(莱斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Xuan Liu, Jingbin Qian, Zhanyu Liu, Hefeng Zhou

中文总结 AI 辅助

针对短视频观看时长预测仅输出标量估计的局限,提出分布式服务接口(DSI),通过分布提供器、低维摘要和轻量读取器,在三个数据集上MAE最低,并优于九种基线。

中文摘要 AI 辅助

观看时长是短视频信息流中的主要参与信号,其预测直接影响排序和曝光。现有方法通过纠正时长偏差或建模更丰富的分布来改进观看时长预测,但大多数在服务时仅暴露期望观看时长或去偏后的观看时长。即使视频时长可供后续模型使用,该接口也仅提供观看时长的单一估计,而不提供完成、过度播放或其他与下游任务相关区域的概率。为解决这一局限,我们提出了分布式服务接口(DSI),其包含一个分布提供器、一个紧凑的低维摘要,以及针对每项任务定制的轻量级读取器。提供器学习一个基于观看比率及其事件时间的四个观看状态的联合分布;基于视频时长的规则去除不兼容的组合,而恢复损失保持秒级精度。摘要将该分布缩减为一小组事件概率、相对于时长的尺度以及不确定性统计。训练提供器后,我们固定其参数,并训练结合摘要与原始上下文的值和排序读取器。在KuaiRec、KuaiRand-1K和WeChat21上,完整的DSI系统在所有三个数据集上实现了最低的平均绝对误差(MAE),比九个基线中最强的结果提高了1.9%至8.5%,并在两个数据集上取得了最佳的XAUC。在比较完整系统时,它还引领了考虑视频时长的检索指标。在匹配读取器保持不变的情况下,摘要保留了超出预测均值与视频时长配对的相关任务信息。使用相同的轻量级线性头处理每个新目标,它还在两个新的观看时长目标上表现最佳,并改善了一个单独记录的参与目标,而随机初始化的提供器无法复现这一增益。

英文摘要

Watch time is the primary engagement signal in short video feeds, and its prediction directly affects ranking and exposure. Existing methods improve watch time prediction by correcting duration bias or modeling richer distributions, but most expose only an expected or debiased watch time at serving time. Even when video duration is available to later models, the interface gives only one estimate of watch time and no probabilities for completion, overplay, or other regions relevant to downstream tasks. To address this limitation, we propose the Distributional Serving Interface (DSI), which has a distribution provider, a compact, low-dimensional summary, and lightweight readouts tailored to each task. The provider learns a joint distribution over four watch states derived from watch ratio and their event times; rules based on video duration remove incompatible combinations, while a restoration loss preserves accuracy in seconds. The summary reduces this distribution to a small set of event probabilities, time scales relative to duration, and uncertainty statistics. After training the provider, we fix its parameters and train value and ranking readouts that combine the summary with raw context. Across KuaiRec, KuaiRand-1K, and WeChat21, the complete DSI system achieves the lowest MAE on all three datasets, beating the strongest result among nine baselines by 1.9% to 8.5%, and achieves the best XAUC on two. It also leads retrieval metrics that account for video duration when complete systems are compared. With matched readouts held constant, the summary retains information relevant to each task beyond a predicted mean paired with video duration. Using the same lightweight linear heads for each new target, it also performs best on two new watch-time targets and improves a separately logged engagement target, while a randomly initialized provider does not reproduce this gain.

↑