arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12253cs.CLcs.AIcs.LG

单个冻结模拟器不够:多智能体强化学习中的模拟器崩溃问题

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi

首次发表
浏览论文内容

中文总结 AI 辅助

针对人机交互多智能体强化学习中单个LLM模拟器导致的策略泛化缺陷,提出Verbalized Sampling和Co-Training两种方案,在多轮基准测试和真实用户研究中显著提升了性能,发布了开源框架SCOPE。

中文摘要 AI 辅助

用于人机交互的多智能体强化学习通常依赖单个大语言模型(LLM)来模拟用户行为。我们表明这种方法在泛化性上存在系统性缺陷,并将该缺陷归因于模拟器崩溃:由于模拟器LLM存在模式崩溃,针对其训练的LLM策略会过拟合到利用模拟器主导模式的狭窄策略,这类策略难以迁移到未见过的模拟器和真实用户。我们从理论上对这种崩溃进行了形式化,并提出两种互补的解决方案,分别用于推理阶段和训练阶段。推理阶段的解决方案是Verbalized Sampling,该方法通过从语言化响应分布中采样来拓宽模拟器的行为,减少模式崩溃;训练阶段的解决方案是Co-Training,该方法针对一组可训练的模拟器联合优化策略,防止其过拟合到任何单个模拟器的模式。我们在三个多轮基准测试Persuasion for Good、$\tau^2$-bench和CooperBench上验证了这两种解决方案。Verbalized Sampling相比单模拟器RL将保留的成功率提升了高达9%,Co-Training则进一步将提升幅度推至14%;人类研究在真实用户上也显示出类似的提升。两种方案均保留了在单模拟器RL下会崩溃的策略多样性。为支持该方向的进一步研究,我们发布了SCOPE,这是一个用于多智能体RL的群体协同训练的开源框架。更广泛地说,我们的结果表明,训练环境的多样性(而非仅策略的多样性)对于多轮RL向真实世界部署的泛化至关重要。

英文摘要

Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theoretically and propose two complementary solutions, one at inference time and one at training time. The inference-time solution, Verbalized Sampling, broadens the simulator's behavior by sampling from a verbalized response distribution, reducing mode collapse. The training-time solution, Co-Training, jointly optimizes the policy against a population of trainable simulators, preventing it from overfitting to any single simulator's mode. We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $τ^2$-bench, and CooperBench. Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users. Both solutions preserve the policy diversity that collapses under single-simulator RL. To support further work in this direction, we release SCOPE, an open-source framework for Population Co-Training multi-agent RL. More broadly, our results suggest that the diversity of the training environment, not only the policy, is critical to the generalization of multi-turn RL to real-world deployment.

发表机构

  • Northeastern University(东北大学)
  • New York University(纽约大学)
  • UC Berkeley(加州大学伯克利分校)
  • University of Washington(华盛顿大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑