AI 中文总结
该研究针对工业推荐流水线的缺陷,提出基于智能体方法的DREAM架构,通过意图引擎、元引擎与奖励双循环优化,在淘宝测试中实现多项核心指标提升,验证了智能体元控制的可行性。
AI 中文摘要
工业推荐系统通常采用级联的召回、排序和重排流水线。尽管这些流水线效率较高,但它们将信息和目标分散到各模块,依赖僵化的规则,对实时意图的感知有限,未能充分处理会话级别的浏览、对比和购买行为的转变。我们提出DREAM(Developing Recommender Engine with Agentic Methods,基于智能体方法开发推荐引擎),这是一种自主优化控制架构,在不替换现有流水线的前提下,于其上方添加了一层可感知、可编排且可审计的策略层。DREAM包含两个核心组件:其一,三层意图引擎(Intent Engine)将设备端信号融合为结构化的L0/L1/L2意图表示,其边缘-云端触发链将报告量降低至约8.7%;其二,元引擎(Meta Engine)使用元模型(MetaModel)进行M1到M2到M3的分层推理,即意图摘要、基于策略记忆(Strategy Memory)的策略规划及参数转换,并通过带有安全护栏的统一出口分发所得参数。奖励双循环(Reward Dual Loop)通过结合离线仿真以探索策略空间、在线反馈以校准结果,持续优化这两个组件,形成生成、执行、评估和经验积累的循环。在淘宝首页信息流上进行的大规模A/B测试显示,仅重排控制一项就使IPV提升2.06%、核心IPV提升2.39%、GMV提升0.88%;将控制扩展至精细排序后,这些增益分别提升至2.71%、3.06%和1.31%,同时PV始终提升超过1%。这些增益无需替换流水线模型,也不影响服务稳定性,证明智能体元控制是工业推荐的可行范式。
英文摘要
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.
CommentsTechnical Report