arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SynCo:基于多智能体强化学习的自进化大语言模型的数据合成协同训练

SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning

Wei Yang, Shawn Li, Yuehan Qin, Yawei Wang, Mingxi Wang, Shixuan Li, Tiankai Yang, Jiate Li, Jesse Thomason, Xuezhe Ma, Yue Zhao

arXiv 2610.11345首次发表:更新:

发表机构

University of Southern California; The George Washington University; Georgia Institute of Technology(南加州大学; 乔治华盛顿大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SynCo是基于多智能体强化学习的自进化LLM数据合成协同训练框架,联合优化合成器与推理器,在8个数学推理基准上性能优于现有方法,增益多来自未解决问题。

AI 中文摘要

自进化大语言模型(LLM)智能体有望通过持续交互与学习实现自主改进,减少对人工整理监督的依赖。要实现这一目标,不仅需要更新智能体本身,还需随着其能力变化进化训练经验。然而,现有大多数流程依赖静态数据集或单独更新的合成模型,导致此前有用的任务变得琐碎,而过于困难的任务仍无信息价值。智能体能力与训练经验之间日益扩大的不匹配限制了持续自我改进。为解决该问题,我们提出SynCo——一种基于多智能体强化学习的自进化LLM的智能体数据合成协同训练框架。SynCo联合优化两个独立参数化的智能体:根据推理器(Reasoner)的进化能力状态构建训练任务的合成器(Synthesizer),以及从所得经验中学习的推理器。每个合成任务会引发多次推理器滚动,其结果为两个智能体提供互补奖励:正确性反馈改进推理器,而任务质量、答案可靠性和基于结果的可教学性则指导合成器。它们的更新会反馈到后续合成轮次,使任务解决策略及其训练分布共同进化。在8个数学推理基准上开展的大量实验表明,SynCo显著优于多种现有合成数据方法和受控基线,实现了最强的整体性能,且其大部分增益来自此前未解决的问题。

英文摘要

Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning. SynCo jointly optimizes two independently parameterized agents: a Synthesizer that constructs training tasks from the Reasoner's evolving capability state, and a Reasoner that learns from the resulting experience. Each synthesized task induces multiple Reasoner rollouts whose outcomes provide complementary rewards to both agents. Correctness feedback improves the Reasoner, while task quality, answer reliability, and outcome-grounded teachability guide the Synthesizer. Their updates are fed back into subsequent synthesis rounds, allowing the task-solving policy and its training distribution to evolve together. Extensive experiments across eight mathematical reasoning benchmarks demonstrate that SynCo substantially outperforms a broad range of existing synthetic-data methods and controlled baselines, achieving the strongest overall performance while deriving most of its gains from previously unsolved problems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑