热带强化学习
Tropical Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对期望回报求和无法支持组合推理的问题,提出热带强化学习,通过取最大值替代求和,实现可复用路径的拼接,并在四个智能体任务上比最强基线高16个百分点。
中文摘要 AI 辅助
针对大型语言模型的强化学习通常最大化期望回报,即对所有成功轨迹的概率求和。然而,经典的求和公式只能报告模型策略成功的频率,而无法指出哪个解决方案实际有效;并且由于概率总和为1,强化一个解决方案可能导致模型遗忘另一个从未被证明是错误的解决方案。这使得期望回报不适合组合推理,在组合推理中,解决方案必须由模型在单独且经常失败的尝试中产生、但很少同时产生的推理步骤组装而成。为解决此问题,我们提出热带强化学习,其基于一个简单的代数变换:不是将替代解决方案的概率相加,而是取它们的最大值,从而得到热带半环。于是,状态的价值变为其最可能验证解决方案的对数概率,并附带一个可重放和复用的显式路径。这实现了真正的组合,因为共享状态处相遇的最佳前缀和最佳后缀即使来自不同的轨迹也可以拼接。为将这一方法付诸实践,我们引入了TROPIC,一种针对确定性、可重置且结果可验证环境的训练算法。在四个智能体任务(Sokoban、Countdown、FrozenLake、WebShop)上,TROPIC比最强的在策略基线高出多达16个百分点。因此,改变强化学习的代数(而不仅仅是其估计器)可以显著提升语言模型中的组合推理能力。
英文摘要
Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models
发表机构
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
- TUM(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。