AI 中文总结
研究针对大语言模型训练困境,引入Skill Self-Play协同进化框架,由提议者、求解器和控制器构成,经强化学习循环使组件协同进化,有效弥合验证与探索差距,推动模型性能提升。
AI 中文摘要
大语言模型训练正从人工设计和标注转向交互驱动的自我进化。现有自我进化方法在任务多样性和验证可靠性间面临根本困境。我们将智能体技能视为调和这一矛盾的有力中间地带,引入技能自我博弈(Skill-SP)框架,它由提议者、求解器和动态技能控制器组成,通过强化学习循环协同进化。实证评估表明Skill-SP能推动模型性能提升。
英文摘要
LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fundamental dilemma between task diversity and verification reliability: environment-bound methods obtain precise feedback but confine learning to narrow domains, while open-ended self-generation broadens the task space but lacks reliable verification, allowing misleading rewards to pollute the training loop. We identify agent skills as a powerful middle ground to reconcile this tension: each skill ensures deep, verifiable execution in a specific scenario, while dynamic routing across skills maintains open-ended task variety. Leveraging this insight, we introduce Skill Self-Play (Skill-SP), a co-evolutionary framework comprising a proposer, a solver, and a dynamic skill controller. Orchestrated via a reinforcement learning loop, these components co-evolve in a continuous self-play loop: the proposer generates challenging tasks conditioned on dynamically sampled skills; the solver explores candidate solutions to push its capability boundaries; and the skill controller collects execution feedback to update and expand the skill library. This interactive co-evolution effectively bridges the gap between structured verification and open-ended exploration. Empirical evaluations on tool-use and reasoning benchmarks demonstrate that Skill-SP, serving as a robust evolution engine, consistently pushes the performance ceiling of competent backbones while catalyzing striking turnarounds for initially misaligned models. Our code is available at https://github.com/Qwen-Applications/skill-self-play.