发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TIGPO为长视野LLM智能体提出跨策略更新的持久转换图策略,通过分配固定回滚预算稳定优势估计,在ALFWorld和WebShop上优于现有相关方法。
AI 中文摘要
基于图的策略优化通过将回滚轨迹组织为状态转换图,提升了长视野大语言模型(LLM)智能体的信用分配效果。然而,现有方法在每次策略更新内独立构建图,丢弃了早期策略发现的转换,且将优势估计限制在小型的批次局部回滚组中。我们提出时间实例图策略优化(Temporal Instance-Graph Policy Optimization,TIGPO),该方法跨策略更新扩展了基于图的信用分配。TIGPO为每个任务维护一个持久的转换图,允许不同策略版本发现的有效转换共同决定当前回滚的信用。为了主动将当前探索与历史经验重新关联,TIGPO在普通任务采样的探索槽位和对先前探索任务进行延迟重试的重访槽位之间分配固定的回滚预算。对于每次重访,TIGPO将当前回滚组与其对应的早期探索组配对,以构建跨时间参考。扩大后的参考旨在在小型回滚组下稳定相对优势估计,而同一任务上的比较可直接捕捉训练阶段间的策略改进。历史转换和分数仅作为结构和独立的统计参考,绝不会在策略损失中重放。在ALFWorld和WebShop上的实验表明,TIGPO始终优于现有的基于组和基于图的策略优化方法。
英文摘要
Graph-based policy optimization improves credit assignment for long-horizon LLM agents by organizing rollout trajectories into state-transition graphs. However, existing methods construct graphs independently within each policy update, discarding transitions discovered by earlier policies and limiting advantage estimation to small, batch-local rollout groups. We propose \emph{Temporal Instance-Graph Policy Optimization} (TIGPO), which extends graph-based credit assignment across policy updates. TIGPO maintains a persistent transition graph for each task, allowing valid transitions discovered by different policy versions to jointly determine credit for current rollouts. To actively reconnect current exploration with historical experience, TIGPO allocates a fixed rollout budget between Exploration slots for ordinary task sampling and Revisit slots for delayed reattempts of previously explored tasks. For each revisit, TIGPO pairs the current rollout group with its corresponding earlier Exploration group to construct a cross-temporal reference. The enlarged reference is designed to stabilize relative advantage estimation under small rollout groups, while comparison on the same task directly captures policy improvement across training stages. Historical transitions and scores serve only as structural and detached statistical references and are never replayed in the policy loss. Experiments on ALFWorld and WebShop demonstrate that TIGPO consistently outperforms prior group-based and graph-based policy optimization methods.