发表机构
Tsinghua University; The Chinese University of Hong Kong(清华大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgenticGen提出奖励引导的智能体框架,将广告视频生成分解为策略选择和草稿生成两阶段,利用在线反馈训练奖励模型,经DPO和GRPO优化,在TikTok广告系统中显著提升CTR、CVR和Advv。
AI 中文摘要
广告视频生成不仅是一项视频合成任务,还是一个以产品为条件的推理问题,其成功与否通过在线业务指标来衡量。最近的视频基础模型能够从多模态条件生成逼真的片段,但它们并未优化产品应如何转化为有效的广告,也未优化如何根据在线业务反馈改进未来的生成。为闭合这一循环,我们提出了AgenticGen,一个奖励引导的智能体框架,将广告视频生成分解为两个可训练推理阶段:策略选择和草稿生成,从而暴露了在线业务反馈可以监督的优化目标。AgenticGen从累积的在线反馈中学习基于性能的奖励,以及一个与人类质量标准对齐的基于评分标准的补充奖励,然后使用它们来监督策略优化。DPO首先将智能体策略推向在线偏好,GRPO随后通过过程和结果奖励进一步优化两个阶段。离线实验验证了奖励模型和连续策略优化。TikTok广告系统中的在线A/B实验表明,经过DPO和GRPO后的AgenticGen在CTR上比SFT基线提高了2.72%,在CVR上提高了2.63%,在Advv上提高了9.61%。
英文摘要
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do not optimize how a product should be transformed into an effective advertisement or how future generation should be improved from online business feedback. To close this loop, we propose AgenticGen, a reward-guided agentic framework that decomposes advertising video generation into two trainable reasoning stages, strategy selection and draft generation, thereby exposing optimization targets that online business feedback can supervise. AgenticGen learns a performance-based reward from accumulated online feedback and a complementary rubric-based reward aligned with human quality standards, then uses them to supervise policy optimization. DPO first moves the agentic policies toward online preferences, and GRPO further refines both stages with process and outcome rewards. Offline experiments validate the reward models and successive policy optimization. Online A/B experiments in the TikTok advertising system show that AgenticGen after DPO and GRPO improves CTR by 2.72%, CVR by 2.63%, and Advv by 9.61% over the SFT baseline.