发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生产LLM工作流跨运行时执行语义耦合问题,提出绑定自适应平台,将工作流定义为数据流图并编译至流式、SWF或Flink基板,无需改代码,验证多种模式且质量无差异,批处理成本降低。
AI 中文摘要
生产级LLM智能体在多种模式下执行工具调用循环、检索链和组合工作流,但执行语义往往与单一运行时耦合。我们在Rufus中遇到了这一可移植性问题,Rufus是一个拥有大型工具目录、服务于数百万亚马逊客户的对话式AI助手。Rufus支持实时服务、异步后台任务以及评估和内容预生成等高吞吐量批处理工作负载。每种模式都有不同的服务级目标,且通常使用独立的运行时。复用流式编排会使异步和批处理工作负载变得阻塞,并阻止使用批处理推理API,而后者在公布价格上提供50%的折扣。我们提出了一种绑定自适应的智能体执行平台,将工作流定义与执行基板分离。开发人员将工作流定义为类型化数据流图,只需定义一次。该平台将图编译为用于实时服务的进程内流式处理、用于异步执行的持久化AWS SWF编排,或用于批处理推理的分布式Apache Flink流处理。无需更改任何工作流代码。LLM推理被表示为可挂起的图节点,其行为取决于基板:在线流式交付、异步持久重试和离线批量提交。我们跨五种编排模式验证了数十个生产智能体配置:单推理RAG、迭代ReAct、组合PreAct、条件路由和多智能体深度研究。在所有三种绑定中,我们发现输出质量没有可检测的差异。批处理执行将每查询推理成本降低至与公布的批处理API定价一致,同时在生产规模下与流式路径并行运行。
英文摘要
Production LLM agents execute tool-calling loops, retrieval chains, and compositional workflows in multiple modes, yet execution semantics are often coupled to one runtime. We encountered this portability problem in Rufus, a conversational AI assistant with a large tool catalog that serves millions of Amazon customers. Rufus supports real-time serving, asynchronous background tasks, and high-volume batch workloads such as evaluation and content pregeneration. Each mode has distinct service-level objectives and typically uses a separate runtime. Reusing streaming orchestration makes asynchronous and batch workloads blocking and prevents use of batch inference APIs, which offer a 50 percent discount at published prices. We present a binding-adaptive agent execution platform that separates workflow definition from execution substrate. Developers define a workflow once as a typed dataflow graph. The platform compiles the graph to in-process streaming for real-time serving, durable AWS SWF orchestration for asynchronous execution, or distributed Apache Flink stream processing for batch inference. No workflow code changes are required. LLM inference is represented as a suspendable graph node whose behavior depends on the substrate: streaming delivery online, durable retry asynchronously, and batched submission offline. We validated dozens of production agent configurations across five orchestration patterns: single-inference RAG, iterative ReAct, compositional PreAct, conditional routing, and multi-agent deep research. Across all three bindings, we found no detectable difference in output quality. Batch execution reduced per-query inference cost in line with published batch API pricing while operating alongside the streaming path at production scale.
Comments6 pages, 1 Figure, 5 Tables