arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RecSys Factory:限制LLM智能体在工业推荐生命周期内的自主决策点

RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle

Dongyang Ao, Kaixiang Fang, Shijie Xu

arXiv 2608.11241首次发表:更新:

发表机构

Tencent; FiT(腾讯; 腾讯FiT)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RecSys Factory 是限制 LLM 智能体自主到决策点的平台,解决工业推荐系统运营的三角困境,在腾讯三业务线部署 78 天,CLI 调度成功率达 78.6%。

AI 中文摘要

将LLM智能体部署到工业推荐系统运营中会暴露出一种三方张力,我们将其定义为自主-确定性-效率的三角困境:通用自主(理解操作人员意图、零样本生成粘合代码)、工业确定性(符合 schema 的特征提取、无崩溃的 A/B 测试、零合规路径幻觉)以及端到端效率,三者中任意两者可被最大化,以牺牲第三者为代价。我们提出 RecSys Factory,这是一个 LLM 智能体平台,已在腾讯三个异构推荐业务线部署 78 天。其设计原则是将自主限制在决策点而非整个流程,通过三种解构方式实现,每种方式解决三角困境的一个顶点:运行时被解构为三个主机发出的事件源(Claude Code Stop 钩子、企业即时通讯 Webhook、工作流调度器 API);平台在等待阶段无长期运行的守护进程,且在 94% 的挂钟时间(用于等待 Spark 或 GPU 作业)内消耗零 CPU。能力被解构为包含 29 个文件的技能生态系统(共 8971 行代码),其每个技能的陷阱表机械编译为包含 400 个条目的 PitfallStore,将自主限制在预提交流程内的有界类型决策表面。该部署覆盖三个业务线,它们具有不相交的标签语义、A/B 层拓扑和操作人员角色;在三个业务线中的两个观察到了入职时的压缩,这作为案例研究观察结果报告,而非概括性主张,且未与平台前的受控基线进行测量。人类通过“人在回路”卡片协议保留在诊断与执行的边界,该协议作为审计跟踪原语(schema 验证、幂等、可重放)部署,来自为期 8 天、16 次运行的试点。在 78 天的窗口内,该平台记录了 1624 次 CLI 工具调度,总成功率为 78.6%。

英文摘要

Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), industrial determinism (schema-conforming feature extraction, non-crashing A/B, zero compliance-path hallucination), and end-to-end efficiency. Any two can be maximized against the third. We present RecSys Factory, an LLM-agent platform deployed for 78 days across three heterogeneous Tencent recommender business lines. The design principle is autonomy at decision points, not over pipelines, made concrete through three deconstructions that each discharge one vertex of the trilemma. Runtime is deconstructed into three host-emitted event sources (Claude Code Stop hooks, corporate-IM webhooks, workflow scheduler APIs): the platform carries no long-running daemon during the wait phase and consumes zero CPU during the 94% of wall-clock spent waiting on Spark or GPU jobs. Capability is deconstructed into a 29-file skill ecosystem (8,971 lines of SKILL.md) whose per-skill pitfall tables mechanically compile into a 400-entry PitfallStore, confining autonomy to bounded typed decision surfaces inside pre-committed pipelines. Deployment spans three business lines with disjoint label semantics, A/B layer topologies, and operator personas; an onboarding-time compression is observed on two of the three and is reported as a case-study observation, not a generalization claim, and not measured against a controlled pre-platform baseline. The human is retained at the diagnostic-versus-execution boundary via a human-in-the-loop card protocol, deployed as an audit-trail primitive (schema-validated, idempotent, replayable) and reported from an 8-day 16-run pilot. Across the 78-day window the platform recorded 1,624 CLI-tool dispatches at a 78.6% aggregate success rate.

Comments21 pages, 6 figures, 9 tables. Reports a 78-day deployment across three heterogeneous industrial recommender business lines (1,624 CLI-tool dispatches). Companion paper: AutoResearch (P3b), which instantiates the same substrate for autonomous research

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑