Harness as a Language: 一个具有最大表达力的极简智能体框架
Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity
- MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
JAZ框架通过极简的invoke原语实现智能体循环,仅凭提示即可在回忆和持续自我改进任务上超越专门系统,同时降低成本。
AI中文摘要:
现代语言模型智能体围绕“智能体循环”构建,其中LLM被置于一个暴露一组工具的环境中,并通过交替进行工具调用和观察其输出来完全控制工作流程。然而,某些工作流程目前需要超出智能体循环本身的额外工程,例如记忆系统和自我改进系统。我们构建了一个LLM智能体框架JAZ,以探索一个几乎仅包含智能体循环本身的最小化harness在多大程度上能够完成这些专门系统所构建的任务。JAZ暴露了一个基于LLM的原语invoke,并提供了一组内置钩子,允许程序员施加约束和监控。推广现有的代码模式智能体循环,invoke是最简单的循环,满足两个定义性属性:(1) LLM可以编写任意可执行代码,其中可以包含递归的invoke;(2) LLM可见的一切——invoke的所有输入以及其与代码环境的交互历史——都是代码环境中的变量。我们从第一性原理出发论证我们的设计,将invoke视为一个语言原语,表示一个每次调用时由LLM在运行时提供实现的函数。为了验证我们核心invoke原语的设计,我们在传统上通过专门外部harness实现的工作流程上评估invoke——仅使用提示,无需手动设计的工具、harness或外部系统(例如内存或文件系统)。在需要超出上下文窗口的回忆能力的长期工作流程中,JAZ invoke在StuLife的回忆密集型部分上以一半的成本比Letta (MemGPT)高出8%。在持续自我改进方面,JAZ invoke在AppWorld上以更低的成本比ACE高出4%。
英文摘要:
Modern language-model agents are built around the agent loop: the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain capabilities such as long-term memory and self-improvement currently require specialized systems beyond the agent loop itself. We built an LLM agent framework, JAZ, to explore the extent to which a minimal harness that is little more than the agent loop itself can accomplish tasks these specialized systems are built for. JAZ exposes a single LLM-based primitive `invoke` and provides a set of built-in hooks that allow the programmer to apply constraints and perform monitoring. Generalizing existing code-mode agent loops, `invoke` is the simplest loop that satisfies two defining properties: (1) the LLM can write arbitrary executable code that can include recursive `invoke`; (2) everything visible to the LLM - all inputs to `invoke` as well as its interaction history with the code environment - are variables in the code environment. We motivate our design from first principles, viewing `invoke` as a language primitive representing a function whose implementation is provided at runtime by an LLM every time it is called. To validate the design of our core `invoke` primitive, we evaluate `invoke` - with only prompting, no manually designed tools, harness, or external systems (e.g., memory or the file system) - on workflows traditionally implemented through specialized harnesses. On long-horizon workflows requiring recall far beyond the context window, JAZ `invoke` outperforms Letta (MemGPT) by 8% at half its cost on the recall-heavy portion of StuLife. On continual self-improvement, JAZ `invoke` outperforms ACE by 4% at a lower cost on AppWorld.