发表机构
Waseda University; Adelaide University(早稻田大学; 阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工具集成智能体自我进化中规划、执行与评估的协调难题,提出UnifiedPlayers协作框架,通过角色特定奖励在GRPO下联合优化,在十二个推理基准上显著超越基线,验证器检测准确率达84.2%。
AI 中文摘要
自我进化方法通过允许使用工具的智能体生成自己的训练数据,减少了对人工标注轨迹的需求。然而,现有方法通常将轨迹生成与评估分离,依赖无法适应新兴失败模式的静态验证器,或依赖可能强化跨轨迹共享错误的自我一致性信号。联合调整规划、执行和评估提供了一种有前景的替代方案,但引入了根本性的协调挑战:每个组件不断改变用于训练其他组件的数据或反馈。我们通过UnifiedPlayers应对这一挑战,这是一个协作框架,包含生成任务的规划玩家、使用Python工具调用生成多轮轨迹的执行玩家,以及构建可执行验证器的评估玩家。我们设计了角色特定的奖励,在GRPO下协调三个玩家朝向共同的学习目标。在两个模型主干和十二个推理基准上,UnifiedPlayers在数学推理和通用推理任务上分别比最强先前基线高出至少3.5%和3.9%。此外,学习到的验证器实现了84.2%的对抗性检测准确率,而其奖励信号在每问题方差上比自我一致性基线高出2.03倍,提供了更具区分性的验证。这些结果突显了专业玩家之间的协作是通往自我增强工具集成智能体的有前景路径。
英文摘要
Self-evolving methods reduce the need for human-annotated trajectories by allowing tool-using agents to generate their own training data. Yet existing methods typically separate trajectory generation from evaluation, relying on static verifiers that cannot adapt to emerging failure modes or self-consistency signals that may reinforce errors shared across trajectories. Jointly adapting planning, execution, and evaluation offers a promising alternative, but introduces a fundamental coordination challenge: each component continuously changes the data or feedback used to train the others. We address this challenge with \textbf{UnifiedPlayers}, a cooperative framework comprising a Planning Player that generates tasks, an Execution Player that produces multi-turn trajectories with Python tool calls, and an Evaluation Player that constructs executable verifiers. We design role-specific rewards that coordinate the three players toward a shared learning objective under GRPO. Across two model backbones and twelve reasoning benchmarks, UnifiedPlayers outperforms the strongest prior baseline by at least 3.5\% on mathematical reasoning and 3.9\% on general reasoning tasks. Moreover, the learned verifier achieves 84.2\% adversarial detection accuracy, while its reward signal exhibits 2.03$\times$ higher per-question variance than a self-consistency baseline, providing more discriminative verifications. These results highlight cooperation among specialized players as a promising path toward self-enhanced tool-integrated agents.