arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在它们能解决之前:从基础模型预测训练后编码智能体的性能

Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

Tan Yu, Alexander Bukharin, Khushi Bhardwaj, Jennifer Williams, Zirui Liu, Jonathan Lingjie Li, Soumye Singhal, Joseph Jennings, Sanjeev Satheesh, Yash Jain, Ashish Vaswani, Venkat Krishna Srinivasan, Matthew Papakipos, Hyunwoo Kim, Jian Zhang, Oleksii Kuchaiev, Markus Kliegl, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jonathan Cohen, Jiantao Jiao

arXiv 2610.10478首次发表:更新:

发表机构

NVIDIA; University of Minnesota – Twin Cities; University of California, Berkeley(英伟达; 明尼苏达大学双城分校; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出三种基于成功轨迹和验证器的筛选方法,无需冷启动即可预测基础模型经后训练后的编码智能体性能,与SWE-bench Verified排名高度一致。

AI 中文摘要

我们如何预测哪个基础检查点值得进行昂贵的智能体后训练?端到端pass@$K$测试成功行为是否已经出现在基础模型的分布中,但它并不适合智能体编码:许多基础检查点无法可靠地生成完成端到端任务所需的格式良好的工具调用。单次或短视界任务通过将多步交互压缩为固定提示和单个补丁来避免这些工具调用失败,但它们回避了我们关心的核心能力:在仓库演变过程中,在多步工具使用中维持连贯状态。为弥合这一差距,我们将成功的后训练智能体轨迹视为基础模型潜力的前瞻信号。重放每条轨迹并在每个代码更改步骤后重新运行测试,可识别出决定性步骤:该步骤的累积补丁将仓库从失败翻转为通过,证明记录的动作在给定先前上下文的情况下解决了任务。受智能体轨迹覆盖原则的启发,我们在该步骤构建了三个筛选器,它们不需要基础检查点从冷启动驱动测试环境:(i) 决定性动作BPB(每字节比特数)衡量认证动作上的概率质量,(ii) 补丁多项选择测试检查检查点在认证动作与同一验证器拒绝的替代方案之间的选择,以及(iii) 前缀条件pass@$K$评估对功能正确生成的支持,并认可测试接受的任何延续。在十对公开基础模型和后训练模型中,所有三个筛选器对队列的排序与后训练SWE-bench Verified pass@$1$高度一致。由于我们的方法只需要基准的成功轨迹及其验证器,它们可用于将未来的智能体编码基准转化为基础模型评估。

英文摘要

How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end. Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single patch, but they sidestep the core capability we care about: maintaining coherent state over many tool-using steps as the repository evolves. To bridge this gap, we treat successful post-trained agent trajectories as a lookahead signal of base-model potential. Replaying each trajectory and rerunning tests after every code-changing step identifies the decisive step: the first step whose cumulative patch flips the repository from failing to passing, certifying that the recorded action solves the task given the prior context. Motivated by a coverage principle for agentic traces, we build three screens at this step that do not require a base checkpoint to drive the harness from a cold start: (i) Decisive-Action BPB (bits per byte) measures the probability mass on the certified action, (ii) Patch MCQ tests the checkpoint's choice between that action and alternatives rejected by the same verifier, and (iii) prefix-conditioned pass@$K$ evaluates support for functionally-correct generations and credits any continuation that the tests accept. Across ten pairs of public base and post-trained models, all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@$1$. As our methods need only a benchmark's successful trajectories and its verifier, they can be applied to turn future agentic coding benchmarks into base-model evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑