训练全过程中小型递归模型的代码生成能力评估
Evaluating Tiny Recursive Models Across Training for Code Generation
浏览论文内容
中文总结 AI 辅助
本研究对比TRM-AR与对照模型,发现递归模型抗过拟合能力优于参数匹配模型,建议递归代码生成模型应在整个训练轨迹上联合评估拟合度与生成表现。
中文摘要 AI 辅助
代码生成日益依赖大型Transformer模型,其性能随规模提升而增强,但这类模型规模庞大、成本高昂,因此对小型模型存在需求,尤其在数据有限的场景中。递归模型通过复用单个模块增加深度,而非堆叠独立层,以此解决上述问题。现有对这类模型的评估通常采用教师强制拟合(基于真实前缀的下一个词元损失)或任务准确率,且仅在单个检查点上进行,然而代码生成是自由运行的,模型会扩展自身的输出。教师强制训练的优势能否在自由运行生成中保留,以及该优势是否在整个训练过程中存在,仍是未解决的问题。为研究这两点,我们在自然语言转Python代码生成任务中,对约2800万参数的自回归小型递归模型(TRM-AR)与参数匹配、深度匹配的对照模型进行对比,在40个训练轮次和3个随机种子下跟踪拟合度与生成表现。递归模型与深度匹配对照模型的拟合度排名发生两次反转,通过验证损失选择各检查点并分析训练轨迹可得到一致的对比结果。在参数相等时,TRM-AR的拟合度、生成质量和泛化能力均优于参数匹配对照模型,其在验证损失差距上弥补了约45%,在生成质量差距上弥补了约57%,但每步计算成本约为参数匹配对照模型的175倍。不过,在有效深度相等时,更大的Transformer在验证最优状态下的拟合度和生成表现更优,表明TRM-AR的优势在于抗过拟合能力,而非更强的性能。这些发现表明,应对递归代码生成模型在整个训练轨迹上的拟合度与生成表现进行联合评估,而非仅在单个检查点上评估。
英文摘要
Code generation increasingly relies on large transformer models, whose capability advances with scale. Yet such a scale is costly, creating demand for small models, especially where data is limited. Recursive models address this by reusing a single block to add depth rather than stacking independent layers. Such models are typically evaluated by teacher-forced fit (next-token loss on ground-truth prefixes) or task accuracy, at a single checkpoint, whereas code is produced by free-running generation, where the model extends its own output. Whether a teacher-forced advantage survives free-running generation, and whether it holds across training, remains open. To study both, we compare a ~28M-parameter autoregressive Tiny Recursive Model (TRM-AR) on natural-language-to-Python code generation against parameter-matched and depth-matched controls, tracking fit and generation across 40 epochs and three seeds. The fit ranking between the recursive model and the depth-matched control reverses twice. Selecting each checkpoint by validation loss and examining the trajectory yields a consistent comparison. At equal parameters, TRM-AR fits, generates, and generalizes better than the parameter-matched control while recovering approximately 45% of the validation-loss gap and 57% of the generation-quality gap between the two controls, at roughly 175 times the per-step cost of the parameter-matched control. However, at equal effective depth, the larger transformer fits and generates better at its validation optimum, suggesting TRM-AR's advantage lies in resistance to overfitting, not greater capability. These findings suggest that recursive code generation models should be evaluated jointly on fit and generation across the training trajectory rather than at a single checkpoint.
发表机构
- Toronto Metropolitan University(多伦多都会大学)
机构由 AI 辅助整理,请以论文原文为准。