arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.05773cs.AI

超越静态评估:构建用于可扩展智能体强化学习的模拟环境

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

发表机构优步人工智能解决方案
查看机构详情
  • Uber AI Solutions(优步人工智能解决方案)

机构由 AI 辅助整理,请以论文原文为准。

Akshay Arora, Ishan Nigam, Ashutosh Aggarwal, Shefali Bansal, Krishna Singh, Sweta Kumari, Nikhil Mittal, Shariq Farhan, Siddarth Malreddy

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型成为自主智能体后传统静态评估失效的问题,引入AgenticAI-Supervisor平台,通过API和UI驱动构建强化学习环境,运用可验证执行结果、多维奖励塑造及严格验证测试,以客户支持智能体案例展示核心能力及未来工作重点。

中文摘要 AI 辅助

随着大语言模型演变成自主智能体,传统静态评估无法捕捉多步决策。我们引入了AgenticAI-Supervisor,这是一个由API和用户界面驱动的强化学习环境,将环境创建与可扩展执行解耦。通过转向可验证的执行结果,该平台生成高保真轨迹并应用多维奖励塑造。关键的是,我们的框架通过严格的内部状态验证和测试减轻了奖励作弊。通过客户支持智能体案例研究展示了模型优化的一致闭环反馈,首次展示了平台的核心能力。未来工作将专注于诸如计算机使用、工具使用、自动“难住”和边缘情况生成等高级功能。

英文摘要

As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples environment creation from scalable execution. By moving to verifiable execution outcomes, the platform generates high-fidelity traces and applies multi-dimensional reward shaping. Critically, our framework mitigates reward hacking through rigorous internal state validation and testing. This work provides a first look at our platform's core capabilities through a Customer Support Agent case study demonstrating a consistent closed-loop feedback for model optimization. Future work will focus on advanced features such as Computer Use, Tool Use, automated "stumping", and edge-case generation.

↑