arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种自触发智能推送推荐系统

A Self-Triggered Agentic Push Recommendation System

Zhao-Yu Zhang, Qingying Chen, Chunyuan Zheng, Jing Zhou, Jian Sun, Siqi Chen, Leiying Chen, Chuan Zhou, Huiyou Jiang, Xin Tao, Haoxuan Li, Zhouchen Lin

arXiv 2608.01949首次发表:更新:

AI 中文总结

本文提出端到端自触发智能推送系统STEPS,含规划、执行、过滤智能体,已在抖音部署,可提升用户活跃天数、降低权限禁用率并减少计算开销。

AI 中文摘要

推送通知是大规模平台上的关键推荐场景,使系统能够主动触达应用外的用户,以提升长期复访率。然而,设计最优推送系统需在严格的系统资源约束下处理“是否推送及何时推送”问题的复杂动作空间。现有解决方案通常分为两类被动范式:离线建模分配推送时间的预规划频率方法,限制了实时适应性;定期轮询系统的固定间隔触发方法,在过高计算开销与最优时机捕获不足间形成两难困境。此外,这类多阶段框架严重受困于局部最优问题。为克服这些局限,本文提出STEPS,一种主动式自触发端到端智能推送推荐系统,已在拥有超10亿用户的抖音全面部署。STEPS将推送推荐重新表述为自触发智能过程,系统不仅决定是否推送,还决定何时再次调用自身,从而形成平衡实时有效性与效率的闭环。具体而言,STEPS包含两个基于决策Transformer的智能体:采用门控序数回归方法规划下一次系统调用的规划智能体,以及基于轨迹奖励决定是否推送的执行智能体。此外,引入轻量过滤智能体,既控制计算开销,又作为防止不合理规划行为的关键保障。在线A/B测试表明,STEPS使用户活跃天数显著提升0.2843%,推送权限禁用率降低1.9089%,同时过滤智能体将计算开销减少79.42%。

英文摘要

Push notification is a critical recommendation scenario on large-scale platforms, allowing the system to proactively reach users outside the application to improve long-term re-engagement. However, designing an optimal push system requires handling a complex action space for the "whether and when" delivery problem under strict system resource constraints. Existing solutions typically fall into two passive paradigms: pre-planned frequency methods that allocate delivery times via offline modeling, limiting real-time adaptability; and fixed-interval triggering methods that periodically poll the system, creating a strict dilemma between excessive computational overhead and diminished optimal timing capture. Furthermore, such multi-stage frameworks severely suffer from local optima. To overcome these limitations, in this paper, we propose STEPS, a proactive, Self-Triggered End-to-end Agentic Push Recommendation System, which is already fully deployed at Douyin with over 1 billion users. STEPS reformulates push recommendation as a self-triggered agentic process in which the system decides not only whether to send a push, but also when to invoke itself again, thereby forming a closed loop that balances real-time effectiveness and efficiency. Specifically, STEPS consists of two decision transformer-based agents: a planning agent that schedules the next system invocation using a gated ordinal regression method, and an execution agent that decides whether to send a push based on trajectory rewards. Furthermore, we introduce a lightweight filtering agent to both control computational overhead and act as a crucial safeguard against unreasonable planning behaviors. Online A/B testing demonstrates that STEPS significantly increases user active days by 0.2843% and reduces the push permission disablement rate by 1.9089%, while the filtering agent reduces computational overhead by 79.42%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑