arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向Web级电子商务的时序感知复购预测:用于多场景杂货推荐的生存模型

Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation

Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt, Yanan Cao, Sinduja Subramaniam, Evren Korpeoglu, Kaushiki Nag, Kannan Achan

arXiv 2608.28393首次发表:更新:

发表机构

Walmart Global Tech(沃尔玛全球科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对电商复购预测的二分类模型栈,提出用生存模型替代,通过AFT模型及校准方法,在杂货电商数据上实现了更优的性能与效率,明确了校准与排名的权衡关系。

AI 中文摘要

电子商务中的复购推荐通常被建模为二分类问题,即“该客户会在W天内购买该商品吗”,该公式需要为每个感兴趣的时间窗口单独训练一个模型。我们用直接预测复购时间的生存模型替代该模型栈,并在某大型杂货电商平台的数百万客户数据上,对超过30种 ablation(消融)配置进行评估。本研究有三项贡献:第一,经验风险分析显示边际风险略有下降(k≈0.9),与“杂货商品自上次购买后间隔越久,复购概率越高(风险递增,k>1)”的普遍直觉不同。对数正态分布(Log-Normal)在边际拟合上表现最佳(R²=0.998),且排名性能最优,尽管威布尔分布(Weibull)在条件残差拟合上最佳,我们详细分析了这一明显差异。第二,单个加速失效时间(AFT)模型替代了三个按时间窗口划分的二分类器,在每个时间窗口的性能与后者相当或更优,同时总树数量减少了约3倍。在生存目标下,特征重要性发生了重新排序:渠道节奏和近期信号的重要性上升,而总频率计数的重要性下降。第三,一个4参数参数校准方法将原始生存累积分布函数(CDF)映射到各时间窗口的概率,且无跨时间窗口的单调性违反。AFT家族的校准质量相差一个数量级:指数AFT(Weibull k=1)的预期校准误差(ECE)约为1e-4,比对数正态分布低约10倍,而排名指标的相对差异在0.3%以内。我们将指数AFT用于需消耗概率的场景,对数正态分布用于纯排名场景,揭示了单一AFT家族内的原则性校准-排名权衡。

英文摘要

Repurchase recommenders in e-commerce are commonly framed as a binary question asking "will this customer buy this item within W days", a formulation that requires a separately trained model for every horizon of interest. We replace this stack with survival models that predict time-to-repurchase directly, and evaluate them on millions of customers from a major grocery e-commerce platform across more than thirty ablation configurations. Our study makes three contributions. First, an empirical hazard analysis reveals a slightly decreasing marginal hazard (k ~ 0.9), differing from the common intuition that grocery items become more likely to be repurchased the longer since the last purchase (increasing hazard, k > 1). Log-Normal achieves the best marginal fit (R^2 = 0.998) and the best ranking, despite Weibull providing the best conditional residual fit, revealing an apparent discrepancy we analyze in detail. Second, a single Accelerated Failure Time (AFT) model replaces three per-horizon binary classifiers, matching or exceeding each at its own horizon while using roughly 3x fewer total trees. Feature importance reshuffles under the survival objective: channel-cadence and recency signals rise while aggregate frequency counts fall. Third, a 4-parameter parametric calibration maps raw survival CDFs to per-horizon probabilities with zero cross-horizon monotonicity violations. Calibration quality varies by an order of magnitude across the AFT family: Exponential AFT (Weibull k=1) achieves expected calibration error (ECE) ~1e-4, roughly 10x lower than Log-Normal, while ranking metrics agree within 0.3% relative. We adopt Exponential AFT for probability-consuming surfaces and Log-Normal for pure ranking, exposing a principled calibration-ranking trade-off within a single AFT family.

CommentsAccepted at ReSys 2026 RecTemp Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑