发表机构
Waseda University; Adelaide University(早稻田大学; 阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自进化工具集成智能体反馈局限,提出AnchorLoop,引入历史执行器冻结副本作为锚点,在13个推理基准上分别提升数学与通用推理2.5%和2.8%。
AI 中文摘要
自进化的工具集成智能体在其自身训练循环内生成的任务和反馈中学习。课程智能体生成任务,而执行智能体通过强化学习从自一致性信号中学习。然而,仅依赖当前执行智能体的反馈存在两个局限:在全共识下,群体相对优势消失;而基于不确定性的课程奖励偏向于分歧,却未显示生成的任务是否支持进一步学习。这些局限促使需要超越当前执行智能体的额外参考。我们提出 AnchorLoop,引入前一轮迭代执行智能体的冻结副本作为历史参考,并在训练循环的两端复用该副本。对于执行智能体,锚点提供交叉参考优势,将当前输出与当前和历史多数答案进行比较。对于课程,它提供基于采样多数一致性差异的协议参考。由于在课程训练期间执行智能体和锚点具有相同参数,此比较作为任务选择的代理,而非跨版本改进或正确性的证据。在13个推理基准上,AnchorLoop在数学推理上比Agent0提高2.5%,在通用推理任务上提高2.8%。它还保持更高的有效优势方差,并在后期迭代中持续改进,而未锚定的基线则收益递减。这些结果证明了在自进化工具集成智能体中引入轻量级历史参考的益处,无需外部任务或答案监督。
英文摘要
Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback has two limitations: group-relative advantages vanish under full consensus, while uncertainty-based curriculum rewards favor disagreement without showing whether the generated tasks support further learning. These limitations motivate an additional reference beyond the current Executor. We propose \textit{AnchorLoop}, which introduces a frozen copy of the previous iteration's Executor as a historical reference and reuses it on both sides of the training loop. For the Executor, the anchor provides a cross-reference advantage that evaluates current outputs against both current and historical majority answers. For the Curriculum, it provides an agreement-based reference based on differences in sampled majority agreement. Since the Executor and anchor have identical parameters during Curriculum training, this comparison serves as a proxy for task selection rather than evidence of inter-version improvement or correctness. Across 13 reasoning benchmarks, AnchorLoop improves over Agent0 by 2.5\% on mathematical reasoning and 2.8\% on general reasoning tasks. It also maintains higher effective-advantage variance and continues improving in later iterations as the unanchored baseline shows diminishing gains. These results demonstrate the benefit of introducing a lightweight historical reference into self-evolving tool-integrated agents without external task or answer supervision.