发表机构
NetEase, Inc.(网易公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对工业推荐系统从研究到上线流程依赖人力的问题,提出AutoLR系统,结合多专家委员会等三种机制,实现从研究到上线的自主管控。
AI 中文摘要
优化工业推荐系统是一个迭代的研究与工程流程,而非从想法直接到部署的路径。在网易游戏社区应用DASHEN中,算法工程师通常从研究论文、技术报告和过往生产实验中识别有前景的方向;复现或适配底层方法;在生产代码库中实现这些方法;并通过训练和离线实验评估生成的模型。有前景的候选方案随后进入在线A/B测试,那些表现出稳定提升的方案会被提交至上线评审——全量流量上线的内部关卡。大型语言模型(LLM)可协助此工作流的各个阶段,但在缺乏能在持续多日的实验周期中可靠协调各阶段的管控工具时,整个流程仍依赖人力。我们提出AutoLR,最初构建为自动上线评审(Auto Launch Review),之后向上扩展为从研究到上线的自主管控工具。AutoLR结合了三种系统机制:一是多专家委员会,对提案进行辩论和对抗性评审;二是确定性证据加权探索-利用选择器,在候选方向间分配有限的试验预算并使用委员会重新排序;三是分层知识系统,结合外部研究、生产系统知识及DASHEN特定的领域知识,如游戏社区、玩家特征和内容交互模式,以及来自配置、补丁、日志、故障和离线结果的后验证据。LLM智能体执行语义推理和代码生成,而确定性控制器保留对执行、指标提取、安全防护及持久状态转换的控制权。
英文摘要
Improving an industrial recommender is an iterative research-and-engineering process rather than a direct path from idea to deployment. In \textbf{DASHEN, NetEase's gaming-community app}, algorithm engineers typically identify promising directions from research papers, technical reports, and prior production experiments; reproduce or adapt the underlying methods; implement them in the production codebase; and evaluate the resulting models through training and offline experiments. Promising candidates are then advanced to online A/B tests, and those demonstrating robust gains are submitted to Launch Review---the internal gate for full-traffic rollout. Large language models (LLMs) can assist with individual stages of this workflow, but the overall process remains human-dependent without a harness that can reliably coordinate them across long-running, often multi-day experimental cycles. We present \textbf{AutoLR}, initially built as \textbf{Auto Launch Review} and later extended upstream into an autonomous research-to-launch harness. AutoLR combines three system mechanisms: a \textbf{multi-expert council} that debates and adversarially reviews proposals; a \textbf{deterministic evidence-weighted exploration--exploitation selector} that allocates a limited trial budget across candidate directions and uses Council reranking; and a layered knowledge system that combines external research, production-system knowledge, and DASHEN-specific domain knowledge---such as game communities, player characteristics, and content-interaction patterns---with posterior evidence from configurations, patches, logs, failures, and offline outcomes. LLM agents perform semantic reasoning and code generation, while deterministic controllers retain authority over execution, metric extraction, guardrails, and persistent state transitions.