从结果到行动:利用事后诸葛亮进行长期语言智能体训练
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
浏览论文内容
中文总结 AI 辅助
研究长期语言智能体训练中现有强化学习方法面临的挑战,提出事后诸葛亮策略优化方法,通过投影到意图空间提取低方差学习信号,聚合相似状态和行动,理论和实证证明可稳定提升策略性能。
中文摘要 AI 辅助
强化学习(RL)已成为改进复杂任务上大语言模型(LLMs)的广泛采用技术。尽管有进展,但现有RL方法在训练长期交互智能体时仍面临挑战,主要瓶颈是区分长期交互中不同行动的贡献,导致高优化方差。为解决此问题,我们引入一种新的策略梯度方法,即事后诸葛亮策略优化(HPO),它将当前策略分布和事后诸葛亮分布投影到意图空间,并从它们之间的瓦瑟斯坦距离中提取低方差学习信号。我们从理论和实证上表明,在意图空间中聚合语义相似的状态和行动可产生有界方差估计器并稳定提高策略性能。我们的代码可在线获取。
英文摘要
Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this progress, existing RL methods still face challenges in training agents with longer-horizon interactions. One major bottleneck is distinguishing the contribution of different actions in long-horizon interaction, leading to high optimization variance. To address this, we introduce a novel policy gradient method, Hindsight Policy Optimization (HPO), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low-variance learning signals from the Wasserstein distance between them. We theoretically and empirically show that aggregating semantically similar states and actions in the intent space yields a bounded-variance estimator and improves policy performance stably. Our code is available online.