arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在回滚结束之前:面向长视界编码代理的早期终端奖励预测

Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents

Jihan Yao, Sihan Zeng, Shangbin Feng, Zhiyuan Fan, Banghua Zhu, Yulia Tsvetkov

arXiv 2609.31995首次发表:更新:

AI 中文总结

针对长视界编码代理奖励稀疏问题,提出上下文早期奖励(CER)方法,通过轨迹前缀行为预测终端奖励,在SWE-bench上提升性能并显著降低推理成本。

AI 中文摘要

长视界编码代理只有在完成昂贵的工具调用序列之后才能获得可验证的奖励。这增加了推理成本,放大了早期的错误假设,并可能导致稀疏的终端奖励和不稳定的训练。我们引入了上下文早期奖励(CER),它通过轨迹前缀中的行为证据来预测终端奖励。CER通过从相关历史任务中总结的经验,综合出针对当前任务和阶段的自适应评分标准。在SWE-bench Verified上的测试时扩展中,CER在Nemotron 3 Ultra上将RM@8相对于最强基线提高了4.2个百分点(pp),在Qwen 3.6 27B上提高了2.0个百分点;在Nemotron上,它仅使用15.3%的令牌即可匹配最佳基线性能。在强化学习训练实验中,CER超过完整回滚TMax 1.9个百分点,同时使用的在线策略和评判令牌减少了52.7%。总之,CER为长视界编码代理提供了一种可解释、高效且密集的评估方法。

英文摘要

Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through behavioral evidence in a trajectory prefix. CER synthesizes adaptive rubrics specific to the current task and stage through experiences summarized from related historical tasks. In test-time scaling on SWE-bench Verified, CER improves RM@8 over the strongest baseline by 4.2 percentage points (pp) on Nemotron 3 Ultra and 2.0 pp on Qwen 3.6 27B; on Nemotron, it takes only 15.3% tokens to match the best baseline performance. In RL training experiments, CER exceeds full-rollout TMax by 1.9 pp while using 52.7% fewer online policy-and-judge tokens. Together, CER provides an interpretable, efficient, and dense evaluation method for long-horizon coding agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑