发表机构
Tote AI(Tote AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LazyAgent通过按需物化与反向闭包,仅执行目标所需步骤,显著降低无关开销,并在生产与基准测试中验证了高效性与等价性。
AI 中文摘要
当前的智能体运行时通常先规划后执行,并且一旦某个步骤就绪便会执行它。我们提出了LazyAgent,一个统一的执行框架,用于围绕一个实时的、由目标派生的需求集合来组织智能体编写的程序。LazyAgent会根据执行状态的变化,从请求的输出中刷新一个反向闭包,并且仅当活动目标需要某个就绪节点时才对其进行物化。这种方法用一次线性时间的图分析以及随后的常数时间成员测试,取代了重复的局部判断,使得程序可以保持广泛性,而执行则保持针对特定请求。在那些描述的内容超出当前请求所需的程序上,LazyAgent通过拒绝在无关工作开始之前就将其排除,持续优于最强的目标停止型急切执行基线。增加一个无关产品会使急切执行的开销增加22.5%,而LazyAgent的开销增加0.0%。LazyAgent在生产科学工作流中节省了42.0%的实测CPU时间,并在一个跨越四个仓库的实时发布门禁中节省了51.7%的容器时间。我们还证明并验证了当请求覆盖整个图时,两者在等价性上是完全一致的,此时没有无关工作可避免。除了权限之外,目标相对的输出投影在两个第三方测试套件上,对一个共享步骤最多可节省约90%的开销,而相同的急切执行控制则节省0.0%;当被省略的输出没有其他消费者或请求需要它时,这种优势会消失。排序、重用和剪枝也可以节省成本,但不能取代权限。最后,我们表明当前的公开基准是急切形状的,并且几乎不包含未请求的工作。一项预注册的规划干预并未使其变得更广泛。这些发现激励了基于常驻程序和序列构建基准。
英文摘要
Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around a live, goal-derived demanded set. LazyAgent refreshes a backward closure from requested outputs as execution state changes and materializes a ready node only when the active goal requires it. This replaces repeated local judgments with one linear-time graph analysis followed by constant-time membership tests, allowing programs to remain broad while execution stays request-specific. On programs that describe more than the current request needs, LazyAgent consistently outperforms the strongest goal-stopping eager baseline by refusing unrelated work before it starts. Adding one unrelated product raises the eager bill by 22.5% and LazyAgent's by 0.0%. LazyAgent saves 42.0% of measured CPU on production scientific workflows and 51.7% of container time on a live release gate spanning four repositories. We also prove and verify exact equivalence when the request reaches the whole graph, leaving no unrelated work to avoid. Beyond permission, goal-relative output projection saves up to approximately 90% of a shared step on two third-party test suites while the identical eager control saves 0.0%; the advantage disappears when the omitted output has no other consumer or the request needs it. Ordering, reuse, and pruning can also save cost, but do not replace permission. Finally, we show that current public benchmarks are eager-shaped and contain almost no unrequested work. A pre-registered planning intervention did not broaden them. These findings motivate benchmarks built from standing programs and sequences.