关键之处信用分配:面向终端智能体的依赖感知策略优化
Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
浏览论文内容
中文总结 AI 辅助
针对终端智能体多步任务中信用分配未显式追踪读写依赖的问题,提出依赖感知组策略优化(DepGPO),通过构建命令依赖图并向后追踪验证器资源来重分配优势,提升复杂终端任务的性能与训练稳定性。
中文摘要 AI 辅助
使用终端的智能体在编码、调试及其他多步骤终端任务中受益于强化学习(RL)。在这些任务中,后续命令通常依赖于先前命令产生的信息或中间结果。然而,现有的轨迹级和步骤级信用分配方法并未显式追踪命令影响最终结果的读写依赖关系。因此,训练信号仍可能被分配给无关操作,削弱了对相关步骤的学习。本文提出依赖感知组策略优化(DepGPO),利用命令之间的执行依赖关系来指导终端智能体的信用分配。具体而言,我们从执行轨迹构建命令依赖图,并从任务验证器检查的资源出发向后追踪。然后,我们将信用分配给相关写入及其沿这些路径的支持读取,并据此在步骤间重新分配轨迹优势。大量对比实验和消融研究表明,DepGPO在复杂终端任务上提升了任务性能和训练稳定性。
英文摘要
Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.
发表机构
- School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。