arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.28019cs.LG

面向开放网络的用户基础模型构建

Building a User Foundation Model for the Open Web

  • Teads(泰兹)

机构由 AI 辅助整理,请以论文原文为准。

Solal Vernier, Ivan Can Arisoy, Merwan Barlier, Blaž Škrlj

AI总结:

该研究针对开放网络RTB场景下用户身份碎片化问题,构建基于自监督学习的用户基础模型,经LLM优化预训练后,在竞价胜率、CTR等生产任务中取得显著性能提升。

AI中文摘要:

用户基础模型在电商和社交推荐领域已展现出优异性能,但多数工业部署假设用户身份稳定且持久。开放网络实时竞价(RTB)的数据分布存在显著差异:用户身份在不同浏览会话中碎片化且不持久,浏览历史的可用性取决于用户隐私选择。因此,大量流量无历史数据,可用记录常由较短、不连贯的会话组成。该领域的历史信号通常表示为聚合计数器和近期区间,未利用序列结构。为解决这一局限,本文提出一种用户基础模型,该模型对用户浏览历史应用自监督学习,证明学习到的表示可提升多个下游生产任务,展示了该方法在开放网络上的可行性。我们预训练了一个Transformer编码器,采用掩码语言建模和序列级对比目标,随后在点击预测任务上进行微调。我们通过对可审查的代码级编辑(提升器)目录进行LLM在环搜索,优化编码器的预训练流水线,在工业场景中实现了LLM作为优化器的范式。相同的编码器表示使生产竞价胜率模型的RIG提升+1.197%,生产CTR排序器的RIG提升+1.354%;为期7天的在线A/B测试确认CTR提升+2.13%,eCPC降低-1.13%(两个指标的80%置信区间均不包含零)。

英文摘要:

User foundation models have demonstrated strong results in e-commerce and social recommendation, but most industrial deployments assume environments where user identity is stable and persistent. Open-web real-time bidding (RTB) operates on a structurally different data distribution: user identity is fragmented and non-persistent across browsing sessions, and the availability of browsing history depends on user privacy choices. Consequently, a significant portion of traffic carries no historical data, and available records often consist of relatively short, disjointed sessions. As a result, historical signals in this domain are typically represented as aggregated counters and recency buckets, leaving the sequential structure unexploited. To address this limitation, we present a user foundation model that applies self-supervised learning on user browsing histories and show that the learned representation improves multiple downstream production tasks, demonstrating the viability of this approach on the open web. We pre-train a Transformer encoder with masked language modeling and a sequence-level contrastive objective, then fine-tune it on the click prediction task. We optimize the encoder's pre-training pipeline with an LLM-in-the-loop search over a curated catalog of reviewable, code-level edits (lifters), instantiating the LLM-as-optimizer paradigm in an industrial setting. The same encoder representation yields +1.197% RIG on the production bid win-rate model and +1.354% RIG on the production CTR ranker; a 7-day live A/B test confirms +2.13% CTR, -1.13% eCPC (80% CI excluding zero on both metrics).

补充信息

↑