AI 中文总结
针对现有RL训练范式的静态对应智能体不匹配问题,提出IB-RL方法,在Vehicle TeleSales和Deal-or-NoDeal任务上较单边RL基线显著提升了策略对未见过对应智能体的泛化能力。
AI 中文摘要
强化学习(RL)在提升大型语言模型(LLM)在具有固定、可验证奖励的任务(如数学推理和代码执行)上的表现方面已取得显著成果。在这些场景中,环境遵循固定规则,不会对智能体进行策略性适应。而策略性对话则不同:环境是另一个会根据策略进行适应的智能体,成功取决于双方的交互。尽管存在这种交互特性,当前的RL方法通常是针对固定的对应智能体或模拟器来训练目标智能体。我们发现这种训练范式会促使策略利用对应智能体特有的规律性,而非学习能泛化到不同对应智能体的策略,我们将此问题称为静态对应智能体不匹配,并在实验中直接对其进行了量化。为解决该问题,我们提出孤立双边强化学习(IB-RL),其中两个角色通过联合rollout共同进化,同时每个角色通过完全独立的优势函数、动作掩码和更新路径来优化自身奖励。我们在两个领域中评估了冻结策略对完全独立的预留对应智能体的表现:在Vehicle TeleSales任务中,IB-RL的Success@1达到89.6%,而最佳的单边RL基线为84.6%;在Deal-or-NoDeal任务中,其与DeepSeek V4 Pro的达成一致率达到98.4%,而最佳的单边基线为86.4%。这些结果表明,通过严格的每个智能体孤立方式联合训练两个角色,能够产生对未见过的对应智能体更具泛化性的策略。
英文摘要
Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as mathematical reasoning and code execution. In these settings, the environment follows fixed rules and does not adapt strategically to the agent. Strategic dialogue differs in this respect: the environment is another agent that adapts to the policy, and success depends on the interaction between the two sides. Despite this interactive nature, current RL approaches typically train a target agent against a fixed counterpart or simulator. We find that this training paradigm encourages the policy to exploit counterpart-specific regularities rather than learn strategies that generalize across counterparts. We call this problem the static-counterpart mismatch, which we quantify directly in our experiments. To address it, we propose Isolated Bilateral Reinforcement Learning (IB-RL), in which the two roles coevolve through joint rollouts while each role optimizes its own reward through fully independent advantages, action masks, and update paths. We evaluate frozen policies against fully independent held-out counterparts in both domains. On Vehicle TeleSales, IB-RL achieves 89.6% Success@1, compared to 84.6% for the best unilateral RL baseline. On Deal-or-NoDeal, it reaches 98.4% agreement against DeepSeek V4 Pro, compared to 86.4% for the best unilateral baseline. These results indicate that jointly training both roles with strict peragent isolation produces policies that generalize more effectively to unseen counterparts.