发表机构
University of Chinese Academy of Sciences; Institute of Neuroscience, Chinese Academy of Sciences; Tongyi Lab, Alibaba Group; University of California, Los Angeles; Nanyang Technological University(中国科学院大学; 中国科学院神经科学研究所; 阿里巴巴集团通义实验室; 加利福尼亚大学洛杉矶分校; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长视界智能体任务中结果奖励稀疏的问题,提出ProCredit方法,通过重新运行验收检查将进度转化为信用,在多个规模上提升任务完成率。
AI 中文摘要
长视界智能体任务要求智能体通过一系列工具调用来修改环境,其成功由最终状态决定。标准方法在任务结束时分配单一的结果奖励,并比较同一任务下采样的轨迹。因此,没有成功轨迹的组别无法产生训练信号,失败的尝试无法根据其接近完成的程度进行区分,而推进任务的回合与仅查询环境的回合获得相同的信用。先前的工作将比较单位从轨迹细化到步骤,或训练奖励模型以提供中间信号:前者仍仅从最终成功中获取信号,后者则通过模型进行估计。我们观察到,决定成功的验收检查同样可以应用于中间状态,因此进度与结果一样可验证。我们提出ProCredit,它将这种已验证的进度转化为信用:它在每个回合后重新运行验收检查,根据进度的变化奖励该回合,并利用这些奖励在相同任务的不同尝试之间以及轨迹内的不同回合之间分配信用。从三个规模的Qwen3.5基础模型开始,在AppWorld上,ProCredit在两个测试集上的每个规模的任务完成率均优于结果奖励基线和基于进度的基线,在4B规模上超过最强结果奖励基线4.1个百分点,且第二个环境中的结果显示了相同的改进方向。消融实验表明,仅将最终进度添加到轨迹分数中并不能提高性能:收益来自于将进度归功于其发生的回合。
英文摘要
Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields no training signal, failed attempts cannot be told apart by how close they came to completion, and turns that advance the task receive the same credit as turns that only query the environment. Prior work refines the unit of comparison from the trajectory to the step, or trains a reward model to supply intermediate signal: the former still derives its signal from final success alone, and the latter estimates it with a model. We observe that the acceptance checks that decide success can also be run on intermediate states, so progress is as verifiable as the outcome. We propose ProCredit, which turns this verified progress into credit: it reruns the acceptance checks after each turn, rewards the turn by its change in progress, and uses these rewards to assign credit both across attempts at the same task and across the turns within a trajectory. Starting from Qwen3.5 base models at three scales on AppWorld, ProCredit outperforms outcome-reward baselines and progress-based baselines in task completion rate at every scale on both test sets, exceeding the strongest outcome-reward baseline by 4.1 percentage points at 4B, and results in a second environment show the same direction of improvement. Ablations show that adding the final progress to the trajectory score alone does not improve performance: the gain comes from crediting progress to the turn where it occurs.