arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

验证器即课程:用于跨家族游戏生成的执行门控自蒸馏

The Verifier is the Curriculum: Precision Sets the Return on Search in Code Self-Distillation

Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou

arXiv 2607.09709首次发表:更新:

发表机构

Institute of Science Tokyo; Zhejiang University; National University of Singapore(东京科学研究所; 浙江大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对游戏生成中代码生成器的优化,提出执行门控自蒸馏方法,通过严格启动过滤器提升跨家族泛化能力,在GameCraft-Bench上实验表明该方法显著提高干净生成率和覆盖率,验证器对模型学习有重要作用。

AI 中文摘要

在针对学习到的评判器对代码生成器进行训练后,可以优化代理特征,从而提高分数但不改进工件。我们研究相反的信号:一个确定性的、无需评判器的、不可博弈的过滤器——即生成的项目在无头引擎下能否干净地启动(严格启动)。在这个门控下,拒绝采样自蒸馏增强了跨家族泛化能力。在GameCraft-Bench(将自然语言摘要映射到完整的Godot项目)上,一个14B模型(Qwen3-14B+LoRA)在严格启动下进行蒸馏,使得在四个未见游戏家族上的干净生成率从每个候选8.8%提高到42.2%,三轮后的K选最佳覆盖率从18/25提高到25/25(黄金上限),每轮都是显著提升(p=0.0019,p<1e-4,p<1e-4)。这种提升并非仅仅来自添加数据:完全匹配的黄金复制控制回归到基础模型以下(5.6%对8.8%,p=0.019),而计数匹配分解将第一轮到第二轮的提升分为质量(+8.8个百分点)和数量(+8.5个百分点)通道。最直接的是,仅交换过滤器重新运行循环——通过99.9%生成的宽松BUILD检查代替启动门控——完全消除了提升(回到基础,p=1e-3与启动门控轮相比),分离出验证器精度而非优化器。第二个不可博弈信号,无头执行基础,在各轮中单调上升,并且在匹配预算下产生比黄金复制更多的有基础的候选(16对5),证实了提升是功能性的,而非启动但为空。游戏生成是一个可验证的试验台,用于说明一个道理:验证器即课程——它所认证的就是模型所学的。

英文摘要

Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We study the opposite signal: a deterministic, judge-free filter that asks only whether a generated project launches cleanly under a headless engine (strict-launch). Under this gate, rejection-sampling self-distillation compounds out-of-family generalization: on GameCraft-Bench a 14B model raises the per-candidate clean-launch rate on four held-out families from 8.8% to 42.2% and coverage at 32 candidates from 84% to 100%, the gold references' own ceiling, beating the supervised model on every one of the 25 held-out tasks. The gate costs one engine invocation per candidate: no reward model, no judge. At a fixed admitted count, what governs the loop is verifier precision. Swapping in a lenient build check alone erases the gain (p=0.0012); a matched gold-duplication control regresses below the supervised model. Under a semantic gate on APPS, dialing fuel precision from 1.0 to 0.25 at fixed candidate count prices that fuel linearly: over 23 training seeds, half-clean fuel returns +3.59 percentage points against the +3.69 a linear rate predicts. Under count-matched rejection-SFT only one direction of verifier error carries a measurable cost. Search obeys the same gate: quadrupling the harvest budget is worth +1.62pp behind a strict gate and nothing distinguishable from zero behind a partial-credit one. Recall is nearly free; search pays only through a precise gate: the verifier is the curriculum.

Comments15 pages, 8 figures, 6 tables. v2: substantially revised and extended (new title, new APPS experiments on verifier precision, unbiased coverage estimator, three training seeds)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑