低延迟系统中的工具制作与自我进化大语言模型智能体
Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究如何解决生产LLM智能体因重复生成代码浪费延迟和可靠性的问题,提出用智能工具制作管道取代推理时编码循环,部署该方法使系统更快、更可靠、易操作,降低延迟和错误率,提高可审计性。
AI中文摘要:
生产环境中的大语言模型智能体每次请求时都会为相同的程序步骤重新生成代码,从而浪费延迟和可靠性。我们用一种智能工具制作管道取代这种推理时的编码循环,该管道在部署前将重复的标准操作程序步骤编译成经过验证的、有版本的工具。工具制作器在实时环境中进行综合,收集执行跟踪、观察后端模式和值、生成候选工具并根据标记案例进行修复。在运行时,生产智能体直接调用这些工具,仅在需要时才回退到代码生成。我们将该方法部署在一个履行中心警报分类系统中,在生产中,工具调用将50%延迟降低了42%,在1500个历史警报上,通过抑制重复步骤中的运行间差异,将端到端错误率降低了高达53%。版本化工具还提高了可审计性,并暴露了规范差距和上游数据漂移。我们的结果表明,自我进化的智能体可以使工业大语言模型系统更快、更可靠且更易于操作。
英文摘要:
Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SOP steps into validated, versioned tools before deployment. The tool-maker grounds synthesis in the live environment as it collects execution traces, observes backend schemas and values, generates candidate tools, and repairs them against labeled cases. At runtime, the production agent calls these tools directly and falls back to code generation only when needed. We deploy the approach in a Fulfillment Center alarm-triage system, where an agent diagnoses alarms against a 44-node SOP over heterogeneous metric backends. In production, tool calls reduce p50 latency by 42%. On 1,500 historical alarms, they reduce end-to-end error rate by up to 53% by suppressing run-to-run variance in repeated steps. Because tools return compact structured verdicts, they also enable a simpler direct-call architecture, reducing p50 latency by a further 62% in a controlled ablation. Versioned tools also improve auditability and expose specification gaps and upstream data drift. Our results show that agents that build and maintain their own tool libraries, a key element of self-evolving agents, can make industrial LLM systems faster, more reliable, and easier to operate.