AI 中文总结
本文提出PinSieve,即面向企业内容质量分类的选择性VLM服务智能体,结合受控内存飞轮维护机制,提升审核效率、降低成本并改善信号交付,具有任务可迁移性。
AI 中文摘要
生产环境中的企业AI智能体通常需要具备边界性、有状态性、可观测性和可管控性,而非完全自主。本文介绍PinSieve,这是一个大规模内容质量流水线中的生产案例研究,其部署组件是选择性视觉语言模型(VLM)服务智能体,仅处理上游轻量模型未解决的灰色区域切片,在线输出标量路由分数,并保留受控的人工介入机制。在该切片上,部署的系统过滤的非可操作项是此前生产模块的2.05倍,同时略微降低了估计漏检率;推广后,它将审核效率提升25.7%,归一化运营成本降低16.2%,并将信号交付从次日提升至当日。随后,本文研究了选择性反馈下的受控内存飞轮维护机制,其中介入项默认进行审核,自动通过项主要通过审计采样标注。反馈内存记录路由轨迹、观测路径、审计倾向和重放元数据,用于评估和调试。数据策划智能体在具有代表性、不确定性、时效性和新审核的重放数据上使用有界的提案-验证循环,在批量接受前设置正例率和分数区间的护栏。在六个月生产数据的链式月度刷新中,该设计将平均FNR@50%从代表性随机重放下的17.73%降至13.29%。推理审核智能体验证教师生成的理由,并支持保留/修复/丢弃决策。生产级成果仅归因于已部署的服务智能体;重放和理由审核结果为离线或采样管控的证据。该相同的服务智能体方案已被应用于多个额外的内部信号,显示出超出单个任务的可迁移性。
英文摘要
Enterprise AI agents in production often need to be bounded, stateful, observable, and governable rather than fully autonomous. We present PinSieve, a production case study in a large-scale content-quality pipeline. Its deployed component is a selective vision-language-model (VLM) Serving Agent that operates only on the grey-zone slice left unresolved by lightweight upstream models, exposes a scalar routing score online, and preserves controlled human escalation. On this slice, the deployed system filters 2.05x more non-actionable items than the previous production module while slightly reducing estimated miss rate; after promotion, it improves review productivity by 25.7%, reduces normalized operating cost by 16.2%, and moves signal delivery from next-day to same-day. We then study maintenance through a governed memory flywheel under selective feedback, where escalated items are reviewed by default and auto-passed items are labeled mainly through audit sampling. Feedback Memory records routing traces, observation paths, audit propensities, and replay metadata for evaluation and debugging. The Data Curation Agent uses a bounded proposal-verifier loop over representative, uncertainty, recency, and fresh-review replay, with positive-rate and score-bin guardrails before batch acceptance. In chained monthly refresh over six months of production data, this design reduces average FNR@50% from 17.73% under representative random replay to 13.29%. A Reasoning Review Agent audits teacher-generated rationales and supports keep/repair/drop decisions. Production claims are attributed only to the deployed Serving Agent; replay and rationale-review results are offline or sampled-governance evidence. The same serving-agent recipe has been adopted to several additional internal signals, suggesting transferability beyond one task.
CommentsAccepted at the KDD 2026 workshop "Enterprise AI Agents: From Prototypes to Production."