arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

再试一次,不要回头:小型代码模型中盲重采样优于自我修复

Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models

Yuvraj Verma

arXiv 2607.26117首次发表:更新:

发表机构

Independent Researcher, India

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对MBPP+数据集在三种规模代码模型上发现,盲重采样在7B以下表现最优,成本远低于其他重试方法,而基于自身失败尝试的自我修复存在锚定效应导致的性能损失。

AI 中文摘要

自我修复是代码智能体的标准组成部分,它会将失败的程序连同测试输出返回给模型以请求修正,其评估几乎总是以完全不重试的基线为参照。我们认为这种比较混淆了反馈的价值与额外尝试的价值。我们在MBPP+上采用安慰剂对照设计,针对1.5B、3B、7B三种模型规模,比较四种匹配预算的重试条件:盲重采样、无内容的失败通知、真实执行反馈以及带有口头自我反思的反馈。盲重采样是7B以下模型的最强条件,在7B规模下仍与最佳条件统计持平,但消耗的token量减少2.5至5.5倍;在1.5B规模下,基于模型自身失败尝试的条件会导致性能下降6.1个点(p=0.006),而执行反馈的信息内容相比安慰剂没有可测量的增益。我们将此归因于锚定效应:当看到自身之前的尝试时,模型在33%-68%的重试中会生成几乎相同的程序,而盲重采样下该比例仅为2%-14%。两项进一步实验限定了该效应的范围:检索其他任务的解决方案不会产生任何变化(波动范围在±3.5个点内),这将危害限定为自我条件作用而非上下文长度;而反思是唯一可测量地削弱锚定效应的条件,但在成本上仍处于劣势。复制实验排除了两种竞争性解释:该惩罚在全精度下保持不变,且在独立模型家族上可复现。在涵盖两个家族和两种精度的六种配置中,其幅度仅由基线质量预测(相关系数r=0.96)——锚定的代价是对首次错误尝试的固化。

英文摘要

Self-repair - returning a failed program to the model together with its test output and asking for a correction - is a standard component of code agents, and is almost always evaluated against a baseline that does not retry at all. We argue that this comparison confounds the value of the feedback with the value of the extra attempt. Using a placebo-controlled design on MBPP+ at three model scales (1.5B, 3B, 7B), we compare four matched-budget retry conditions: blind resampling, a content-free failure notice, genuine execution feedback, and feedback augmented with verbal self-reflection. Blind resampling is the strongest condition below 7B, and remains statistically tied with the best condition at 7B, while consuming 2.5-5.5x fewer tokens; conditioning on the model's own failed attempt costs 6.1 points at 1.5B (p=0.006), and the informational content of execution feedback adds nothing measurable over the placebo. We attribute this to anchoring: when shown its previous attempt, a model reproduces a near-identical program in 33-68% of retries, against 2-14% under blind resampling. Two further experiments delimit the effect. Retrieved solutions to other tasks change nothing (bounded to +/-3.5 points), which localizes the harm to self-conditioning rather than context length; and reflection, the only condition that measurably weakens the anchor, remains dominated on cost. Replication rules out two competing explanations: the penalty is unchanged at full precision, and it reproduces on an independent model family. Across six configurations spanning two families and two precisions, its magnitude is predicted by baseline quality alone (r=0.96) - the cost of anchoring is the cost of committing to a bad first attempt.

CommentsCode, pre-registrations and run traces: https://github.com/vermayuvraj/self-improving-agent

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑