发表机构
Meta AI(Meta AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出原生大语言模型工具CORAL,将智能体闭环应用于生产级推荐系统,通过A/B实验证实其可在两个大规模社交平台上实现参与度提升或服务成本降低,自动化持续优化工作。
AI 中文摘要
生产级推荐系统决定着数十亿用户的所见内容,要维持其性能就需要持续优化:随着内容、用户行为和上游模型的变化,必须重新审视控制检索、排序和服务的各项选择。传统上,人类工程师通过在线实验测试这类变更——这一过程缓慢且被动,受工程工作量限制,在条件变化时系统的部分内容会得不到修订。尽管大语言模型已被应用于排序、用户建模和离线模型开发,但很少有系统能将智能体置于持续闭环中,使其作用于实时推荐系统并从其决策的实测效果中学习。我们提出了CORAL(Constraint-Optimized Recommender via an Agentic Loop,基于智能体循环的约束优化推荐系统),这是一种原生大语言模型工具,可实现该闭环:每个循环中,智能体观测运行信号,对过往决策和结果的记忆进行推理,并调用工具——包括将变更控制在固定运行预算内的数值优化器——以重新配置推荐系统,实测结果会为下一个循环提供信息。我们将此表述为一个部分可观测、非平稳的约束优化问题,其中策略无需参数更新,而是从其先前的动作中在上下文中改进。在两个大规模社交平台上通过A/B实验评估,该工具在一个平台上无需额外服务成本即可提升参与度,在另一个平台上则可降低服务成本且不降低参与度,覆盖了参与度-效率前沿。性能随循环迭代而提升,表明单一智能体循环可在明确的约束下自动化传统由人类算法工程师执行的持续优化工作。
英文摘要
Production recommender systems shape what billions of people see, and sustaining their performance requires continual optimization: as content, user behavior, and upstream models shift, the choices governing retrieval, ranking, and serving must be revisited. Traditionally, human engineers test such changes through online experiments--a slow, reactive process limited by engineering effort, leaving parts of the system unrevised as conditions change. Although large language models have been applied to ranking, user modeling, and offline model development, few systems place an agent in a continual closed loop that acts on a live recommender and learns from the measured effects of its decisions. We present CORAL (Constraint-Optimized Recommender via an Agentic Loop), an LLM-native harness that closes this loop: each cycle, the agent observes operating signals, reasons over a memory of past decisions and outcomes, and invokes tools--including a numerical optimizer that keeps changes within a fixed operating budget--to reconfigure the recommender, with measured outcomes informing the next cycle. We formulate this as a partially observed, non-stationary, constrained optimization problem in which the policy improves in context, without parameter updates, from its prior actions. Across two large-scale social platforms, evaluated with A/B experiments, the same harness improves engagement at no additional serving cost on one and reduces serving cost without degrading engagement on the other, spanning the engagement-efficiency frontier. Performance improves as the loop iterates, suggesting that a single agentic loop can automate continual optimization work traditionally performed by human algorithm engineers under explicit guardrails.
CommentsAccepted by RecSys '26 OARS Workshop