arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AURA:面向大规模生产推荐系统的智能体诊断与优化

AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale

SungGeun Kim, Abhinav Narain, Daniel Nemirovsky

arXiv 2609.16625首次发表:更新:

发表机构

The Walt Disney Company; Intuit(华特迪士尼公司; Intuit公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对生产推荐系统难以从聚合指标识别用户失败模式的问题,提出端到端智能体系统AURA,通过读取大规模日志进行定性诊断并生成代码级改进,已在流媒体平台验证并具跨领域迁移性。

AI 中文摘要

推荐系统为何以及如何辜负其服务的用户?通常,实践者只能依靠利益相关者团队的反馈、领域专业知识以及数据分析的见解来改进算法。然而,推荐对终端用户表现优劣的具体细节和模式,难以从聚合的定量指标中辨别。这些指标提供的是高层面且不完整的图景,而要深入理解推荐质量及其模式,则需要在大规模范围内结合领域理解与客观性进行推理。我们思考了这一复杂难题,并描述了一种利用最新AI智能体进展的方法和实现,为生产推荐系统提供可操作的诊断和改进。我们提出AURA(推荐算法的智能体理解与优化),一个端到端的智能体系统,能够在大规模下进行定性评估,并在代码层面生成对算法的改进。专门的智能体读取生产参与日志,涵盖数千到数百万个会话,揭示推荐系统如何辜负真实用户的模式和实例。下一步,利用这些诊断以及关于推荐系统自身代码、数据和训练流程的上下文,提出并实施基于该代码库的改进。我们报告了系统设计、在大型媒体流媒体公司两个大型消费平台上的生产数据初步测试、安全措施、运营经验,以及迈向自我改进推荐系统的早期成果。最后,AURA所弥合的诊断差距并非流媒体所特有。该架构旨在可迁移:每个领域特定元素都通过配置层进入,该配置层已将其在两个平台之间移植。我们将其具体映射到电子商务和在线零售推荐场景。

英文摘要

How and why does a recommender system fail the users it serves? Oftentimes, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain expertise, and insights from data analyses. Yet the nuances of how and where recommendations perform well or poorly for end users are difficult to discern from aggregate quantitative metrics. Whereas these metrics provide a high-level and incomplete picture, further granularity into the quality of recommendations and their patterns requires reasoning with domain understanding and objectivity, at scale. We contemplate this complex conundrum and describe a method and implementation that uses the latest AI agentic advances to provide actionable diagnoses and improvements for production recommender systems. We present AURA (Agentic Understanding and Refinement of recommender Algorithms), an end-to-end agentic system that performs qualitative evaluation at scale and can then generate improvements to our algorithms at the code level. Specialized agents read production engagement logs, from thousands of sessions to millions, and surface patterns and examples of how the recommender fails real users. The next step uses those diagnoses and context about the recommender's own code, data, and training pipeline to propose and implement refinements grounded in that codebase. We report the system design, initial tests on production data from two large consumer platforms at a major media-streaming company, safeguards, operational learnings, and early results toward a self-improving recommender system. Finally, the diagnostic gap AURA closes is not specific to streaming. The architecture is built to transfer: every domain-specific element enters through the configuration layer that already ported it between our two platforms. We map it concretely to e-commerce and online-retail recommendation.

Comments14 pages, 1 figure, 6 tables. Accepted at GenAIECommerce'26: The Third Workshop on Agentic and Generative AI for E-Commerce, co-located with RecSys 2026, September 28, 2026, Minneapolis, MN, USA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑