arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08570cs.AI

FailForge:从持续失败中提取过程能力到代码智能体

FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents

Dongyi Lv, Fushun E, Aichen Cai, Liang Huang, Ya Zhang, Qiuyu Ding, Canhui Wu, Zhi Wang, Yuesong Zhang, Jiaqi Wang, Nan Duan

AI总结:

FailForge 框架将代码智能体的失败案例转化为训练信号,以边际成本恢复超26%失败实例,使 Qwen3.5-4B 在 SWE-bench Verified 上的解决率提升6.6个百分点,增益集中于最难问题。

AI中文摘要:

拒绝采样微调(RFT)被广泛用于训练代码智能体,方法是在可验证的软件工程任务上生成轨迹,保留通过测试的轨迹,并对成功的 rollout 进行微调。然而,即使是强大的代码智能体在这类任务的很大一部分上也会反复失败,而标准 RFT 会直接丢弃这些失败案例。被丢弃的样本恰恰是最难且最具信息价值的,它们来自于整理成本高昂的可验证实例。更强的基础模型可能会减少失败的数量,但剩余的困难案例仍然定义了进一步改进的前沿。我们提出 FailForge,这是一个智能体框架,可将失败的 rollout 转化为训练信号。对于每个失败的实例,智能体会从错误反馈和执行轨迹中诊断失败,将诊断提炼为简洁且可操作的技能,并将该技能注入智能体上下文以进行有指导的第二次尝试。在技能指导下成功的轨迹会被重新纳入 RFT 语料库。关键在于,在训练时会移除该技能,因此模型会内化恢复的行为,而不是在推理时依赖外部提示。FailForge 以边际额外成本恢复了超过 26% 的先前失败实例,在增强后的语料库上训练 Qwen3.5-4B,与强大的 RFT 基线相比,SWE-bench Verified 的解决率提高了 6.6 个百分点,且增益集中在最难的问题上。

英文摘要:

Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining those that pass the tests, and fine-tuning on the successful rollouts. However, even strong code agents repeatedly fail on a substantial fraction of such tasks, and standard RFT simply discards these failures. The discarded samples are precisely the hardest and most informative ones, drawn from verifiable instances that are costly to curate. Stronger base models may reduce the number of failures, but the remaining hard cases still define the frontier for further improvement. We propose FailForge, an agentic framework that converts failed rollouts into training signal. For each failed instance, an agent diagnoses the failure from error feedback and execution traces, distills the diagnosis into a concise and actionable skill, and injects the skill into the agent context for a guided second attempt. Trajectories that succeed under skill guidance are folded back into the RFT corpus. Crucially, the skill is removed at training time, so the model internalizes the recovered behavior rather than relying on external hints at inference. FailForge recovers over 26% of previously failed instances at marginal additional cost, and training Qwen3.5-4B on the augmented corpus improves the SWE-bench Verified resolve rate by 6.6 points over a strong RFT baseline, with gains concentrated on the hardest problems.

↑