Co-RL:多智能体强化学习中多样群体涌现的无监督推理
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
浏览论文内容
中文总结 AI 辅助
本研究提出Co-RL框架,通过参数不共享的多智能体协作训练,利用同伴奖励减少对真实标注的依赖,提升纯文本及多模态任务的推理性能,缓解训练崩溃与响应同质化问题。
中文摘要 AI 辅助
强化学习(RL)已成为提升语言及视觉-语言模型推理能力的强大方法,但其最成功的应用仍严重依赖真实标注(如可验证的奖励)。这类标注获取成本高昂,且随着推理能力超出人类可可靠评估的范围,其稀缺性日益加剧。自奖励式RL通过让模型从自身生成的完成内容中推导奖励信号,减少了对真实标注的依赖。然而,仅基于自身生成的反馈进行训练,会强化现有偏差与次优行为,降低响应多样性,最终导致响应同质化与训练崩溃。本研究表明,无监督推理可通过协作多智能体训练涌现。我们提出Co-RL框架,该框架包含多个参数不共享的解耦模型,通过使用来自同伴的奖励信号进行RL同步优化。我们进一步证明,通过异构模型家族、模型规模及改写训练样本增加群体多样性,可减少驱动自增强反馈循环的相关误差。这种多样性持续提升推理性能,维持行为多样性,并缓解训练崩溃。在纯文本与多模态领域,Co-RL在未使用任何真实标注的情况下,始终优于基础模型与现有无标签方法,且表现与监督方法相当或更优。具体而言,Co-RL在7个针对大语言模型(LLM)的纯文本基准测试中,平均提升3.0-8.6%;在4个针对视觉-语言模型(VLM)的多模态基准测试中,平均提升2.3-7.2%。代码可在该https URL获取。
英文摘要
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self-rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions. However, training solely on self-generated feedback can reinforce existing biases and suboptimal behaviors, reduce response diversity, and ultimately lead to homogenized responses and training collapse. In this work, we show that unsupervised reasoning can emerge through cooperative multi-agent training. We introduce Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers. We further show that increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self-reinforcing feedback loops. This diversity consistently improves reasoning performance, maintains behavioral diversity, and mitigates training collapse. Across text-only and multimodal domains, Co-RL consistently outperforms the base models and prior label-free approaches, while matching or surpassing supervised methods, without access to any ground-truth labels. Concretely, Co-RL yields average gains of 3.0-8.6% across seven text-only benchmarks for LLMs and 2.3-7.2% across four multimodal benchmarks for VLMs. Code is available at https://github.com/DrStranded/Co-RL.
发表机构
- University of Exeter(埃克塞特大学)
- ByteDance(字节跳动)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- UC San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。