TAGR:面向工业级直播广告的时间自适应生成式推荐
TAGR: Temporally Adaptive Generative Recommendation for Industrial Live-Streaming Advertising
浏览论文内容
中文总结 AI 辅助
本文提出TAGR框架,通过时间自适应解决现有生成式推荐器在直播广告场景的不足,在电商平台部署后显著提升了直播间进入率、购物车点击率及营收。
中文摘要 AI 辅助
直播广告是短视频和电商平台重要的变现渠道,其中快速变化的直播内容、推广商品及用户反馈对推荐模型提出了极强的时效性要求。现有针对静态领域设计的生成式推荐器在三个层面存在不足:静态语义ID(SID)无法跟踪动态变化的直播广告;单尺度行为建模会遗漏用户意图的转移;新鲜的在线策略反馈与训练稳定性之间存在偏好优化冲突。本文提出TAGR,一种在三个层面具备时间自适应能力的生成式推荐框架:直播广告分词、用户意图建模及偏好对齐。在分词层面,直播语义协作ID(LSID)会基于每个活跃广告当前的直播场景和推广商品定期更新其SID,同时保留稳定的分层分词词汇表以支持自回归生成。在意图层面,意图感知生成(IAG)将用户的直播间进入历史按多个时间粒度建模为主要意图序列,将辅助行为作为单独输入,并使用请求后的意图证据和业务价值对下一个分词预测(NTP)进行加权。在对齐层面,间歇性在线策略偏好优化(IOPO)会定期从当前策略中采样新鲜候选组,执行与行为和价值对齐的偏好更新,并穿插监督式NTP维护以保留学习到的行为分布。在一个大规模电商直播广告平台部署后,TAGR使直播间进入率和购物车点击率分别提升8.5%和7.4%,相比生产基线实现16.1%的营收提升。这些结果证明了面向直播广告的时间自适应生成式推荐的有效性和工业可行性。
英文摘要
Live-streaming advertising is an important monetization channel on short-video and e-commerce platforms, where rapidly changing live content, promoted products, and user feedback impose strong freshness requirements on recommendation models. Existing generative recommenders designed for static domains fail at three levels: static semantic IDs (SID) cannot track evolving live ads; single-scale behavior modeling misses shifting intent; preference optimization conflicts between fresh on-policy feedback and training stability. We propose TAGR, a generative recommendation framework with temporal adaptation at three levels: live-ad tokenization, user intent modeling, and preference alignment. At the token level, Live Semantic-Collaborative ID (LSID) periodically refreshes each active ad's SID based on its current live scene and promoted products, while retaining a stable hierarchical token vocabulary for autoregressive generation. At the intent level, Intent-Aware Generation (IAG) models live-room entry histories at multiple temporal granularities as the primary intent sequence, keeps auxiliary behaviors as separate inputs, and weights next-token prediction (NTP) using post-request intent evidence and business value. At the alignment level, Intermittent On-Policy Preference Optimization (IOPO) periodically samples fresh candidate groups from the current policy and performs behavior- and value-aligned preference updates interleaved with supervised NTP maintenance to preserve learned behavior distribution. Deployed on a large-scale e-commerce live-stream advertising platform, TAGR improves live-room entry and shopping-cart click rates by 8.5% and 7.4%, respectively, and achieves a 16.1% revenue lift over the production baseline. These results demonstrate the effectiveness and industrial viability of temporally adaptive generative recommendation for live-stream advertising.