arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Agent行为即代码:具有程序化规范的高效且稳健的LLM智能体

Agent Behavior as Code: Efficient and Robust LLM Agents with Programmatic Specifications

Peng Qi, Chunliang Lyu, Gang Li, Fabian Chan, Cheng Chang, Ignacio Cases, Will Lu

arXiv 2610.04824首次发表:更新:

发表机构

Uniphore(Uniphore)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ABCAgent,用符号程序完全指定智能体行为,由FM编辑程序保持灵活性,在多个基准上超越神经智能体,显著提升稳健性和效率。

AI 中文摘要

基于基础模型(FM)的AI智能体已展现出执行复杂开放式任务的强大能力。然而,它们在实践中面临一些常见挑战:(a)即使对于语义相似的任务,智能体行为也可能发生剧烈偏差,导致灾难性的错误传播;(b)由于FM调用,每当任务以不同输入重复出现时,都会产生高成本和延迟;(c)FM有限的上下文和指令遵循能力限制了智能体管理不断增长的执行上下文和遵循复杂计划的能力。我们提出了Agent行为即代码(ABCAgent),它使用符号程序(例如,带有潜在神经函数的Python代码)在运行时完全指定智能体的行为,并由一个强大的FM智能体编辑该程序以实现灵活性。因此,行为在无需过早绑定变量的情况下被指定,并且其执行是确定性的。我们在六个智能体基准上评估了ABCAgent,其中两个是我们构建的,用于测试派生程序对其编写任务变体的泛化能力。ABCAgent在GAIA和增强版GAIA上匹配了模型匹配的神经智能体,并在稳健性和长控制流重要的场景中超越它:在GSM-Symbolic上达到98.3%对比97.3%(p=0.001),在τ²-bench的电信领域上达到71.9%对比47.4%的Pass^4(p=0.0001),并在我们控制流增强的WorkArena基准上,在每个循环长度下都记录了更多正确写入的记录。对于更参数化的任务族,ABCAgent在效率上也显著更优。无需编写新程序,ABCAgent解决了GSM-Symbolic实例的92.6%和增强版GAIA变体的20.1%,这使得在GSM-Symbolic上延迟降低5.2倍,成本降低7.0倍,在增强版GAIA上成本降低19%,在τ²-telecom上智能体延迟降低9.5倍。

英文摘要

AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice: (a) agent behavior can deviate drastically even for semantically similar tasks, leading to catastrophically propagated errors; (b) high cost and latency due to FM calls, repeated in full whenever a task recurs with different inputs; (c) FMs' limited context and instruction following capability confine how well agents manage the ever-growing execution context and follow complex plans. We introduce $\textbf{A}$gent $\textbf{B}$ehavior as $\textbf{C}$ode $\textbf{Agent}$ (ABCAgent), which uses a symbolic program (e.g., Python code with potential neural functions) to fully specify the agent's behavior at runtime, with a powerful FM agent editing that program for flexibility. Behavior is thus specified without premature variable binding, and its execution is deterministic. We evaluate ABCAgent on six agent benchmarks, two of which we construct to test how well a derived program generalizes to variants of the task it was written for. ABCAgent matches a model-matched neural agent on GAIA and augmented GAIA, and surpasses it where robustness and long control flows matter: 98.3% against 97.3% on GSM-Symbolic ($p = 0.001$), 71.9% against 47.4% $\mathrm{Pass}^4$ on the telecom domain of $τ^2$-bench ($p = 0.0001$), and more records written correctly at every loop length on our control-flow-augmented WorkArena benchmark. For more parametric task families, ABCAgent is also significantly superior in efficiency. Without authoring a new program, ABCAgent solves 92.6% of GSM-Symbolic instances and 20.1% of augmented GAIA variants, which yields $5.2\times$ lower latency and $7.0\times$ lower cost on GSM-Symbolic, 19% lower cost on augmented GAIA, and $9.5\times$ lower agent latency on $τ^2$-telecom.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑