DCEO:面向电商搜索中长期用户价值建模的直接因果效应优化
DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search
浏览论文内容
中文总结 AI 辅助
本文针对电商搜索中长期用户价值建模的粒度对齐问题,提出DCEO框架,通过演员-评论家结构优化相对因果效应,在线A/B测试中GMV较传统方案提升0.36%。
中文摘要 AI 辅助
工业级电商搜索系统的最终目标是优化用户层面的长期指标,如n天累计购买量或每位用户的商品交易总额(GMV)。然而,这类目标定义在用户层面,而搜索排序则基于每次请求内的商品层面评分。现有方法通常通过人工设计的多目标融合来弥合这种粒度差距,将点击、加购、购买及交易价值等多个商品层面目标的预测结果组合成排序评分,作为最终目标的代理。这类人工设计的融合方案依赖于少量手动调整的权重,限制了细粒度个性化,导致与最终目标的对齐效果不佳。本文提出DCEO(Direct Causal Effect Optimization,直接因果效应优化),一种用于学习与最终目标更对齐的商品层面代理评分的数据驱动框架。我们首先将商品层面代理评分聚合为用户层面代理指标,并使用相对因果效应量化其与最终目标的对齐程度。随后开发了一种演员-评论家框架:评论家估计给定用户层面代理指标的最终目标,演员则动态生成多个目标的上下文依赖融合权重以构建商品层面代理评分,并接受训练以直接优化相对因果效应。大量离线实验与分析证明了DCEO的有效性与可解释性。此外,DCEO已部署至大规模工业级电商搜索系统,在为期41天的在线A/B测试中,其GMV表现较传统GMV代理方案提升了0.36%。
英文摘要
Industrial e-commerce search systems ultimately aim to optimize the user-level long-term objective, such as n-day cumulative purchases or gross merchandise value (GMV) per user. However, such objectives are defined at the user level, whereas search ranking is based on item-level scores within each request. Existing methods typically bridge this granularity gap through manually designed multi-objective fusion, where predictions of multiple item-level objectives, such as clicks, carts, purchases, and transaction value, are combined into a ranking score that serves as a proxy for the ultimate objective. Such hand-crafted fusion schemes rely on a small set of manually tuned weights, limiting fine-grained personalization and leading to suboptimal alignment with the ultimate objective. In this paper, we propose DCEO (Direct Causal Effect Optimization), a data-driven framework for learning item-level proxy scores that are better aligned with the ultimate objective. We first aggregate the item-level proxy scores into a user-level proxy metric and quantify its alignment with the ultimate objective using a relative causal effect. We then develop an actor-critic framework, where the critic estimates the ultimate objective for a given user-level proxy metric, and the actor dynamically generates context-dependent fusion weights over multiple objectives to construct the item-level proxy scores and is trained to directly optimize the relative causal effect. Extensive offline experiments and analyses demonstrate the effectiveness and interpretability of DCEO. In addition, DCEO has been deployed in a large-scale industrial e-commerce search system, outperforming the conventional GMV proxy by 0.36% in GMV in a 41-day online A/B test.
发表机构
- Taobao & Tmall Group of Alibaba(阿里巴巴淘宝天猫集团)
机构由 AI 辅助整理,请以论文原文为准。