arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DREAM技术报告

DREAM Technical Report

Bin Zhang, Bowen Zheng, Chao Yi, Chengyu Lai, Dian Chen, Dimin Wang, Gaoyang Guo, Jialin Zhu, Jian Wu, Jing Yu, Jiuning Lin, Lingqing Zhang, Lingyun Zheng, Mao Zhang, Mingming Pan, Ruiquan Lan, Shuai Zhong, Wen Chen, Wendong Zhang, Xiaodong Zhu, Xuan Chen, Xunke Xi, Yifan Lu, Yiheng Wang, Yue Zeng, Yujie Luo, Yuning Jiang, Zhe Hu, Zhibo Xiao, Zihong Huang, Binbin Cao, Bo Zheng, Danning Wang, Dixuan Wang, Ge Fan, Haixia Wu, Han Zhu, Hao Fang, Haoming Chen, Huiping Chu, Jian Wang, Jianjun Wu, Jiawei Wu, Jiaxin Yu, Jingwen Liu, Jinzhe Shan, Kai Meng, Kai Zhang, Keqin Xu, Kewei Zhu, Lang Tian, Leihui Chen, Li Chen, Licheng Xu, Lide Xiao, Ruitong Zhang, Shiyao Peng, Silu Zhou, Tao Wang, Wei Shi, Wenjun Yang, Xiang Chen, Xiang Gao, Xiao Ren, Xu Liu, Xuwen Wang, Yang Li, Yeqiu Yang, Yi Hu, Yichen Yuan, Yinnan Song, Yipeng Yu, Yuan Liu, Yunqi Gao, Zhiliang Huang, Zhujin Gao, Zongyuan Wu

arXiv 2608.09408首次发表:更新:

AI 中文总结

该研究针对工业推荐流水线的缺陷,提出基于智能体方法的DREAM架构,通过意图引擎、元引擎与奖励双循环优化,在淘宝测试中实现多项核心指标提升,验证了智能体元控制的可行性。

AI 中文摘要

工业推荐系统通常采用级联的召回、排序和重排流水线。尽管这些流水线效率较高,但它们将信息和目标分散到各模块,依赖僵化的规则,对实时意图的感知有限,未能充分处理会话级别的浏览、对比和购买行为的转变。我们提出DREAM(Developing Recommender Engine with Agentic Methods,基于智能体方法开发推荐引擎),这是一种自主优化控制架构,在不替换现有流水线的前提下,于其上方添加了一层可感知、可编排且可审计的策略层。DREAM包含两个核心组件:其一,三层意图引擎(Intent Engine)将设备端信号融合为结构化的L0/L1/L2意图表示,其边缘-云端触发链将报告量降低至约8.7%;其二,元引擎(Meta Engine)使用元模型(MetaModel)进行M1到M2到M3的分层推理,即意图摘要、基于策略记忆(Strategy Memory)的策略规划及参数转换,并通过带有安全护栏的统一出口分发所得参数。奖励双循环(Reward Dual Loop)通过结合离线仿真以探索策略空间、在线反馈以校准结果,持续优化这两个组件,形成生成、执行、评估和经验积累的循环。在淘宝首页信息流上进行的大规模A/B测试显示,仅重排控制一项就使IPV提升2.06%、核心IPV提升2.39%、GMV提升0.88%;将控制扩展至精细排序后,这些增益分别提升至2.71%、3.06%和1.31%,同时PV始终提升超过1%。这些增益无需替换流水线模型,也不影响服务稳定性,证明智能体元控制是工业推荐的可行范式。

英文摘要

Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.

CommentsTechnical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑