发表机构
NVIDIA-labs(英伟达实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出NOOA这一Python框架构建可靠AI智能体,采用智能体即Python对象的编程模型,结合六个面向模型的理念,证明当前模型能有效使用此接口并在相关测试中表现良好,为智能体开发提供新途径。
AI 中文摘要
传统智能体开发分散在提示模板、工具模式、回调代码和工作流图中。我们提出了NVIDIA面向对象智能体(NOOA),这是一个用于构建可靠人工智能智能体的与模型无关的Python框架。NOOA采用更简单的方法:智能体是一个Python对象。其方法是模型可采取的动作,字段是其状态,文档字符串是其提示,类型注释是契约。代码体由“...”组成的方法在运行时由LLM驱动的智能体循环完成,而具有常规代码体的方法仍是标准的确定性Python。这为开发者和智能体提供了相同的接口,智能体行为可像其他软件一样进行测试、追踪、重构和改进。本文有三个贡献。一是提出智能体即Python对象的编程模型及其背后的设计原则,通过简单Python API暴露特定智能体功能。二是识别出六个面向模型的理念并首次在单个框架中结合。三是证明当前模型能有效使用此接口,在目标能力测试及相关基准测试中表现良好。
英文摘要
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions the model can take, fields are its state, docstrings are its prompts, and its type annotations are contracts. A method whose code body consists of "..." is completed at runtime by an LLM-driven agent loop, while methods with normal bodies remain standard deterministic Python. This gives developers and agents the same interface, so agent behavior can be tested, traced, refactored, and improved just like other software. This paper makes three contributions. (1) We present the agent-as-a-Python-object programming model and the design principles behind it. Where Python has existing abstractions, we adopt them directly. Agent-specific capabilities--context, events, state rendering, long-term memory, and validated LLM loops--are exposed through simple Pythonic APIs, so both developers and agents share one familiar programming model. (2) We identify six model-facing ideas that NOOA is, to our knowledge, the first to combine on a single surface: typed input/output, pass-by-reference over live objects, code as action, programmable loop engineering, explicit object state, and model-callable harness APIs for context and events. We find the community already converging on several of these ideas--often as experimental or partial features--and present the comparison to encourage further adoption. (3) We demonstrate that current models use this interface effectively, both in targeted capability tests and on agentic and reasoning benchmarks such as SWE-bench Verified and Terminal-Bench 2.0 and ARC-AGI-3.