arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从可行失败前缀中学习:面向长程LLM智能体的里程碑可行性潜力策略优化

Learning from Viable Failure Prefixes: Milestone Viability Potential Policy Optimization for Long-Horizon LLM Agents

Qi Zhou, Yuanfan Li

arXiv 2609.37111首次发表:更新:

AI 中文总结

针对长程LLM智能体在稀疏奖励下信用分配失效的问题,提出MVPO算法,从可行失败前缀中学习,在ALFWorld和WebShop上分别提升+4.4和+5.3成功点。

AI 中文摘要

长程LLM智能体需要强化学习方法,以便在稀疏和延迟奖励下为中间决策分配信用。现有的基于组的方法(如GRPO和GiGPO)通过比较轨迹回报或重复的锚定状态来缓解这一问题,但当比较的回报没有变化时,它们仍然会失败。我们将这种失败模式识别为零信用失败:在早期训练中,许多失败的轨迹包含有用的前缀,但现有方法未赋予它们任何任务判别性优势。为解决此问题,我们提出里程碑可行性潜力策略优化(MVPO),一种从可行失败前缀中学习的潜力路由策略优化算法。MVPO在并查集可行性区域上估计前缀潜力,用潜力差优势修复零信用组,并根据相对性能进展衰减潜力分支。使用Qwen2.5-1.5B-Instruct的实验表明,MVPO优于包括GRPO和GiGPO在内的八个强基线。在相同训练长度下,MVPO在ALFWorld上比GiGPO基线提高+4.4个成功点,在WebShop上提高+5.3个成功点,同时仅增加0.16%-0.20%的优势构建开销。

英文摘要

Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Existing group-based methods such as GRPO and GiGPO alleviate this issue by comparing rollout returns or repeated anchor states, but they still fail when the compared returns have no variation. We identify this failure mode as zero-credit failure: during early training, many failed rollouts contain useful prefixes, yet existing methods assign them no task-discriminative advantage. To address this issue, we propose Milestone Viability Potential Policy Optimization (MVPO), a potential-routed policy optimization algorithm that learns from viable failure prefixes. MVPO estimates prefix potential over Union-Find viability regions, repairs zero-credit groups with potential-difference advantages, and attenuates the potential branch according to relative performance progress. Experiments with Qwen2.5-1.5B-Instruct show that MVPO outperforms eight strong baselines, including GRPO and GiGPO. Under the same training length, MVPO improves over the GiGPO baseline by +4.4 success points on ALFWorld and +5.3 on WebShop, while adding only 0.16%-0.20% advantage-construction overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑