教智能体可靠地编写代码
Teaching Agents to Code Reliably
浏览论文内容
中文总结 AI 辅助
本文提出通过训练策略而非依赖脚手架来提升自主编码智能体的可靠性,利用执行反馈引导搜索和基于回退树的补丁评分,在SWE-bench Verified上以更少的步骤解决更多问题,并显著提升pass@1和验证器精确率。
中文摘要 AI 辅助
自主编码智能体通过阅读代码、运行命令、编辑文件和提交补丁来解决代码库问题。额外的推理时计算只有在产生有用的修复并为选择提供可靠证据时才能带来收益。三种行为决定了这两点,我们认为它们是可教的,而非规模的副产品,因此策略可以承载它们,而不是依赖脚手架。位置多样性仍然狭窄,因为尝试会回到同一位置,额外的样本不会增加覆盖率。编辑多样性未被利用,因为不同的方法论能解决互补的问题,而单次运行无法达到。验证会误导,因为智能体为其自身补丁编写的测试会接受许多不正确的补丁。通过执行反馈引导搜索,并根据每个补丁自身的回退树对其进行评分,解决了SWE-bench Verified中52.8%的问题,同时使用了八样本基线所花费的48.1%的智能体步骤。训练将这些行为移入策略中。在从SFT和RL训练中保留的270个问题上,加权监督微调将pass@1从31.9%提高到35.2%,pass@8从46.7%提高到51.1%。随后,强化目标针对金标准修复和错误变体训练验证器,并奖励检测到这些错误的断言。它将pass@1提高到43.0%,pass@8提高到60.7%,将验证器精确率从26.8%提高到41.7%,并将错误接受率降低了一半以上。在两个非分布套件中的两个以及所有三个套件中的验证器精确率均有所提高,并且在7B、14B和30B规模下,相对于已发布的编码器基线,这些收益得以保持。
英文摘要
Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three behaviors decide both, and we argue they are teachable rather than byproducts of scale, so a policy can carry them instead of a scaffold. Location diversity remains narrow, since attempts return to the same site and extra samples add no coverage. Edit diversity is left unexploited, since methodologies that differ resolve complementary issues no single run reaches. Verification misleads, since a test the agent writes for its own patch accepts many incorrect ones. Directing search by execution feedback and scoring each patch against its own reverted tree resolves 52.8% of SWE-bench Verified using 48.1% of the agent-steps an eight-sample baseline spends. Training moves these behaviors into the policy. On the 270 issues held out from SFT and RL training, weighted supervised fine-tuning raises pass@1 from 31.9% to 35.2% and pass@8 from 46.7% to 51.1%. A reinforcement objective then trains the verifier against gold-labeled repairs and incorrect variants, crediting the assertions that detect them. It raises pass@1 to 43.0% and pass@8 to 60.7%, lifts verifier precision from 26.8% to 41.7%, and more than halves false acceptance. Resolution improves on two of three out-of-distribution suites and verifier precision on all three, and the gains hold at 7B, 14B, and 30B against published coder baselines.
发表机构
- Stanford University(斯坦福大学)
- AWS AI Labs(AWS AI实验室)
机构由 AI 辅助整理,请以论文原文为准。