arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15591cs.AI

Agent Gym:通过人在回路反馈实现LLM智能体持续评估与演化的框架

Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback

Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit

首次发表
浏览论文内容

中文总结 AI 辅助

Agent Gym是支持LLM智能体持续评估演化的模块化框架,通过人在回路反馈实现行为修正,其发票处理参考验证表明框架实用可用。

中文摘要 AI 辅助

部署在生产环境中的大语言模型(LLM)智能体面临一个根本矛盾:智能体的行为在部署时就已固定,而它必须处理的业务规则和边缘场景却在不断演化。现有方法聚焦于智能体构建和一次性评估,但无法提供在不修改智能体源代码的情况下对部署后行为进行持续修正的结构化机制。市场上多数现有方法需要大量收集日志和轨迹,且需工程团队重新审视智能体设计,这一过程繁重、耗时,抵消了智能体化转型的经济价值。我们提出Agent Gym,这是一个模块化、领域无关的框架,可将任何现有基于LLM的智能体封装进持续评估与演化循环。该框架提供六种可组合能力——Act(执行)、Evaluate(评估)、Investigate(调查)、Correct(修正)、Learn(学习)和Observe(观测),分布在三个架构区域:宪法层(constitution layer),通过配置工件对领域知识进行编码;运行时推理流水线,串联执行、调查和自适应修正;学习循环,使主题专家能通过自然语言交互发现并验证新的修正规则。关键技术贡献包括:具有21个条件算子和三层操作的混合确定性-LLM修正引擎、用于无基准合规性验证的三层调查架构,以及在人工批准前保证规则正确性的程序化安全循环。我们还提出了Spec-to-Note Gap,一种受自编码器启发的智能体系统透明度视图。针对发票处理的开源参考实现表明,该框架完全可运行且已可投入使用。

英文摘要

Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve. Existing approaches address agent construction and one-time evaluation but provide no structured mechanism for continuous post-deployment behavioral correction without modifying the agent's source code. Most of the approaches offered in the market, require intense collection of logs and traces, and re-examining the agent design by the engineering team, a process which is heavy, long and negates the economical value of agentic transformation. We introduce Agent Gym, a modular, domain-agnostic framework that wraps any existing LLM-based agent in a continuous evaluation-and-evolution loop. The framework provides six composable capabilities --- Act, Evaluate, Investigate, Correct, Learn, and Observe --- organized across three architectural zones: a constitution layer that codifies domain knowledge in configuration artifacts, a runtime inference pipeline that chains acting, investigation, and adaptive correction, and a learning loop that enables subject matter experts to discover and validate new correction rules through natural language interaction. The key technical contributions include a hybrid deterministic-LLM correction engine with 21 condition operators and three-tier actions, a three-layer investigation architecture for ground-truth-free compliance validation, and a programmatic safety loop that guarantees rule correctness before human approval. We further introduce the Spec-to-Note Gap, an autoencoder-inspired view of agentic system transparency. An open-source reference implementation for invoice processing demonstrates that the framework is fully operational and ready for adoption.

发表机构

  • Google Cloud(谷歌云)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑