发表机构
Queen’s University; Centre for Software Excellence, Huawei Technologies(女王大学; 华为技术有限公司软件卓越中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从工业视角探讨LLM后训练的棕地维护挑战,提出数据制品编程的工程规范,通过案例研究表明产出优化补丁可显著提升CodeForces和LiveCodeBench的代码生成性能且无显著退化。
AI 中文摘要
工业后训练是一种棕地维护机制:团队继承已部署的检查点,必须在固定计算资源和混合预算下实现针对性改进,同时不导致其他部分性能退化。被维护的产物日益成为数据制品(dataware):其行为由精心设计的后训练混合体调控,通过有界混合补丁而非全新的重新训练进行更新。从一项工业代码生成改进工作中,我们以维护者视角阐述了该工作在实践中为何困难,提炼出三个反复出现的挑战:零和混合设计、以产出(yield)为绑定指标,以及不确定性下的端到端集成,并指出进展更多依赖数据制品编程的工程规范,而非一次性方案。在我们的案例研究中,将教师蒸馏转化为可用训练数据的干预措施,使可接受的监督信号增加了2.84倍,同时使用相同的解决方案教师和每个候选问题的4次解决方案尝试。在主要评估中,经产出优化的补丁使CodeForces的pass@1提升了2.59个百分点(pass@3提升3.11个百分点),使保留的LiveCodeBench v6的pass@1提升了6.11个百分点(pass@3提升8.05个百分点),所有结果在每个条件下从一个固定检查点对每个基准进行16次随机评估中均具有统计显著性,内部AIME和MATH回归套件在容差范围内。
英文摘要
Industrial post-training is a brownfield regime. Teams inherit a deployed checkpoint and must land targeted improvements under fixed compute and mixture budgets without regressing the rest. The maintained artifact is increasingly dataware: behavior governed by a curated post-training mixture, updated via bounded mixture patches rather than clean-slate retraining. From an industrial code-generation improvement effort, we offer a maintainer's perspective on why this work is hard in practice, distilling three recurring challenges, zero-sum mixture design, yield as the binding metric, and end-to-end integration under uncertainty, and arguing that progress depends less on one-off recipes than on an engineering discipline for programming dataware. In our case study, interventions that raised the conversion of teacher distillation into usable training data increased accepted supervision by 2.84 times while using the same solution teacher and four solution attempts per candidate problem. In our primary evaluation, the yield-engineered patch improved CodeForces pass@1 by +2.59 points (+3.11 pass@3) and held-out LiveCodeBench v6 pass@1 by +6.11 (+8.05 pass@3), all statistically significant across 16 stochastic evaluations of each benchmark from one fixed checkpoint per condition, with internal AIME and MATH regression suites within tolerance.