AI 中文总结
本文实证比较了直接生成与迭代差异编辑两种代码模型训练模式,发现直接生成全面更优,并揭示差异模式仅在局部编辑任务上有效,提出“任务局部性”概念。
AI 中文摘要
用于代码编辑的大型语言模型可以在至少两种输出模式下进行训练和部署:直接生成模式,模型一次性输出整个修改后的文件;以及基于迭代差异的生成模式(“步骤”),模型输出一系列局部搜索/替换编辑,逐一应用,直到模型发出完成信号或步骤预算耗尽。基于差异的模式具有吸引力,因为它反映了开发人员编辑代码的方式,并且每轮应生成更少的令牌。我们在共享的 Flutter/Dart 数据集上,以两种模式训练了两个代码模型——一个从零训练的 100M 参数模型(Rainbow-Pony-100M)和一个微调的 Qwen2.5-Coder-0.5B——并评估了由此产生的四个模型,每个模型在约 1,790 个保留任务上进行了评估。直接生成在我们测量的每个指标上均显著优于基于差异的生成——编译/静态分析通过率、每字节比特数、与参考的字符级相似度,以及盲法 LLM 评审员对目标完成度、正确性和代码质量的评分——并且在对任务难度进行匹配 ID 比较控制后,以及在限制为双方都能编译的代码时,这一差距依然存在。随后,我们确定了基于差异的生成确实获胜的条件背后的一个与架构无关的单一机制:它在短距离、空间局部化的编辑上具有竞争力,并且其类别级胜利恰好集中在两个任务类别——重构和错误处理/边界情况修复——这两个类别在我们数据集中具有最低的平均编辑步骤数。我们将此称为任务局部性,并讨论了其对编辑式训练制度何时适合或不适合代码编辑模型的影响。
英文摘要
Large language models used for code editing can be trained and deployed in at least two output regimes: direct generation, where the model emits the entire modified file in one shot, and iterative diff-based generation ("steps"), where the model emits a sequence of localized search/replace edits applied one at a time until it signals completion or a step budget is exhausted. The diff-based regime is attractive because it mirrors how developers edit code and should require far fewer generated tokens per turn. We train two code models - a 100M-parameter model trained from scratch (Rainbow-Pony-100M) and a fine-tuned Qwen2.5-Coder-0.5B - in both regimes on a shared Flutter/Dart dataset, and evaluate all four resulting models on a held-out set of approx 1,790 tasks per model. Direct generation substantially outperforms diff-based generation on every metric we measure - compilation/static-analysis pass rate, bits-per-byte, character-level similarity to the reference, and blinded LLM-judge ratings of goal fulfillment, correctness, and code quality - and the gap persists after controlling for task difficulty via a matched-ID comparison and when restricting to code that compiles on both sides. We then identify a single, architecture-independent mechanism behind the conditions where diff-based generation does win: it is competitive on short, spatially localized edits, and its category-level wins concentrate in exactly the two task categories - refactoring and error-handling/edge-case fixes - with the lowest mean edit-step count in our dataset. We term this task locality and discuss its implications for when an edit-based training regime is and is not the right choice for a code-editing model.
Comments19 pages, 7 figures