arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33618cs.AIcs.CL

ParaAgent:在开放世界工具环境中强化并行行动

ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments

  • Fudan University(复旦大学)
  • University of Edinburgh(爱丁堡大学)
  • Chinese University of Hong Kong(香港中文大学)
  • Vivo AI(vivo人工智能研究院)
  • Huazhong University of Science and Technology(华中科技大学)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

Shengbin Yue, Hongru Wang, Siyuan Wang, Xiaoxin Chen, Wei Chen, Zhongyu Wei

AI总结:

ParaAgent通过结合阶段级探索与执行及行动级并行的结构化循环,在多级优势解耦强化学习下训练,在开放世界工具基准上超越GPT-4.1等基线,实现最佳平均成功率。

AI中文摘要:

语言模型智能体越来越多地部署在开放世界工具环境中,这要求平衡探索未知能力与利用已知能力。现有方法面临性能与效率的权衡:它们要么僵化地解耦探索与执行,要么无协调地交错进行。我们认为关键不在于是否解耦或交错,而在于如何在多个粒度上协调它们。我们引入ParaAct,一种结构化的并行行动循环,结合了阶段级探索⇌执行与行动级并行。为学习此循环,ParaAgent结合多智能体冷启动演示与多级优势解耦下的强化学习,使规划结构显式化,并用步骤级、阶段级和轨迹级奖励进行监督。学习由我们的ToolEnv支持,这是一个基于50,011个真实工具接口的可扩展模拟器。在两个开放世界工具基准上,ParaAgent-4B在所有基线(包括GPT-4.1系统)中取得了最佳平均成功率,在多工具任务上提升最大。行为分析表明,这些收益源于这种行动组织,突显了其对能力强且高效的开放世界智能体的重要性。

英文摘要:

Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We argue that the key lies not in whether to decouple or interleave them, but in how to coordinate them across granularities. We introduce ParaAct, a structured parallel-action loop that combines phase-level Exploration $\rightleftharpoons$ Execution with action-level parallelism. To learn this loop, ParaAgent combines multi-agent cold-start demonstrations with reinforcement learning under multi-level advantage decoupling, making planning structure explicit and supervising it with step-, phase-, and trajectory-level rewards. Learning is supported by our ToolEnv, a scalable simulator grounded in 50,011 realistic tool interfaces. On two open-world tool benchmarks, ParaAgent-4B achieves the best average success among all baselines, including GPT-4.1 systems, with the largest gains on multi-tool tasks. Behavioral analyses show that these gains stem from this action organization, highlighting its importance for capable and efficient open-world agents.

↑