AgentMercury:你的智能体可规模化合成面向业务场景的可验证环境
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
浏览论文内容
中文总结 AI 辅助
本研究提出AgentMercury框架,可规模化合成业务场景的可执行环境,经其训练的智能体策略在企业工作流及多领域基准上表现提升,且环境构建过程可通过微调学习优化。
中文摘要 AI 辅助
智能体通过与环境交互学习行动,但用于训练的环境通常是手动构建或围绕预定义任务与基准合成的。这种以任务为中心的范式难以扩展到反映现实且不断演进的工作流的环境,在这类环境中,多样的任务可从底层世界自然涌现。我们提出AgentMercury,一个用于从高层级业务场景合成可执行环境的规模化框架。AgentMercury并非为特定任务构建环境,而是先实例化一个包含实体、服务、工具、状态以及可执行跨服务不变量的持久世界,后续多样的任务与交互轨迹可从中涌现。我们构建了涵盖14个行业、50个国家的4783个可执行环境,并将其用作强化学习的训练基底。尽管这些环境的生成未针对评估基准,在这类面向业务的环境上训练的策略在企业工作流及涵盖推理、编码、科学计算、工具使用的域外基准上均大幅提升。实验中,Qwen3.5-4B在AgentMercury环境上训练后,EnterpriseOps-GYM的表现从12.3提升至15.7,AIME26的表现从45.9提升至56.0。我们进一步表明,构建过程本身可被学习:在构建轨迹上微调Qwen3.5-35B-A3B,可在保留的业务场景上将可执行世界创作成功率从3.3%提升至83.3%。这些结果表明,基于场景的环境可提供超越特定基准训练的有用且可泛化的学习信号,而其构建本身可成为一种可学习的能力。
英文摘要
Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environments that reflect realistic and evolving workflows where diverse tasks can naturally emerge from the underlying world. We introduce AgentMercury, a scalable framework for synthesizing executable environments from high-level business scenarios. Rather than constructing an environment for a specific task, AgentMercury first instantiates a persistent world with entities, services, tools, state, and executable cross-service invariants, from which diverse tasks and interaction trajectories can subsequently emerge. We construct 4,783 executable environments spanning 14 industries and 50 countries, and use them as training substrates for reinforcement learning. Despite being generated without targeting the evaluation benchmarks, policies trained on these business-oriented environments improve substantially on both enterprise workflows and out-of-domain benchmarks spanning reasoning, coding, scientific computing, and tool use. In our experiments, Qwen3.5-4B improves from 12.3 to 15.7 on EnterpriseOps-GYM and from 45.9 to 56.0 on AIME26 after training on AgentMercury environments. We further show that the construction process itself can be learned: fine-tuning Qwen3.5-35B-A3B on construction traces increases executable-world authoring success from 3.3% to 83.3% on held-out business scenarios. These results show that scenario-grounded environments can provide useful and generalizable learning signals beyond benchmark-specific training, while their construction can itself become a learnable capability.
发表机构
- Meridian Intelligence Global Inc.(子午线智能全球公司)
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。