发表机构
Meta Platforms, Inc.(元平台公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对在线自动研究中的无效结论问题,提出人工门控的EvoPilot方法,通过角色智能体、版本化技能和持久化验证,在37天视频检索活动中纠正评估缺陷并实现离线3.20个百分点改进和在线0.66%相对提升。
AI 中文摘要
大型语言模型(LLM)智能体能够提出、实现并评估模型变更。自动研究循环通过在自包含程序上进行分钟级迭代展示了这一能力。而在线自动研究则跨越异步系统、数小时级变体和数周级活动,这些活动可能影响实际产品。当代码变更为空操作、数据窗口泄漏、评估器语义漂移或两个实验分支经历不同的服务漏斗时,即使运行完成,仍可能支持无效结论。我们提出EvoPilot,一种用于长周期在线自动研究的人工门控方法。角色特定智能体通过版本化领域技能和类型化适配器执行每一轮操作。持久化记录保留实验和失败;确定性检查强制执行记录的教训。我们针对支撑视频深度探索(VDD)的检索系统进行了一项为期37天的活动,VDD是一种在用户打开种子视频后用于发现后续视频的在线体验。该活动覆盖了七个方向,并使用每小时刷新的数亿视频索引。早期的手动实验未确立交互头带来的收益。一次原始的自动研究尝试重新审视该方向,但错误地将离线命中率下降22个百分点归因于该交互头。随后我们引入了EvoPilot。其人工门控验证将下降追溯到预先存在的评估缺陷,该缺陷产生了3,000和600的输出深度。修复后,匹配比较测得离线改进3.20个百分点。研究后重放和突变测试拒绝了无效比较,同时接受了有效对应项。持久化状态恢复了一轮中断的迭代,工件重用避免了约5个GPU小时。另外,一项为期七天的随机在线评估估计VDD切片中留存良好搜索结果率(GSRR)相对增加0.66%。
英文摘要
Large language model (LLM) agents can propose, implement, and evaluate model changes. Autoresearch loops demonstrate this capability through minutes-scale iterations on a self-contained program. Online autoresearch instead spans asynchronous systems, hours-long variants, and weeks-long campaigns that can influence a product. A completed run can still support an invalid conclusion when a code change is a no-op, data windows leak, evaluator semantics drift, or the two arms traverse different serving funnels. We present EvoPilot, a human-gated method for long-horizon online autoresearch. Role-specific agents execute each round through a versioned domain skill and typed adapter. Durable records preserve experiments and failures; deterministic checks enforce recorded lessons. We study a 37-day campaign for the retrieval system that powers Video Deep Dive (VDD), an online experience for discovering follow-on videos after a user opens a seed video. The campaign covered seven directions and used an hourly refreshed index of hundreds of millions of videos. Earlier manual experiments had not established a benefit from an interaction head. A primitive autoresearch attempt revisited the direction but incorrectly attributed an offline hit-rate decline of 22 percentage points to the head. We then introduced EvoPilot. Its human-gated verification traced the drop to a pre-existing evaluation defect that produced output depths of 3,000 and 600. After repair, a matched comparison measured an offline improvement of 3.20 percentage points. Post-study replay and mutation tests rejected invalid comparisons while admitting valid counterparts. Durable state recovered an interrupted round, and artifact reuse avoided approximately five GPU-hours. Separately, a seven-day randomized online evaluation estimated a 0.66% relative increase in the VDD slice of Good Search Result Rate for Retention (GSRR).
Comments9 pages, 1 figure, 8 tables. ACM sigconf format; submitted to the KDD 2027 Applied Data Science Track