支架介导的后训练:协同演化模型参数与程序性支架图
Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs
浏览论文内容
中文总结 AI 辅助
该研究针对大型语言模型后训练中参数与推理支架脱节的问题,提出支架介导的后训练范式,通过协同演化实现技能提升,在FeatureBench上较标准SFT取得显著性能优势。
中文摘要 AI 辅助
大型语言模型的后训练仅优化参数,而推理时的程序性支架通常独立于参数训练设计,这种脱节导致难以自动获取和内化复杂策略。我们提出支架介导的后训练:程序性支架被组织为可演化图结构,通过发现、蒸馏和动态重编译与模型参数协同演化。我们将该范式实例化为技能训练。在FeatureBench上,自动发现的技能使通过率提升8.1个百分点;经逐步蒸馏后,模型在无任何外部支架的情况下仍达到27.7%的通过率(蒸馏保留率为85.2%,定义为蒸馏后通过率/带技能通过率),显著优于相同数据上的标准SFT。
英文摘要
Post-training of large language models optimizes only parameters, while inference-time procedural scaffolds are typically designed independently of parameter training. This disconnect makes it difficult to automatically acquire and internalize complex strategies. We propose scaffold-mediated post-training: procedural scaffolds are organized into an evolvable graph structure that co-evolves with model parameters through discovery, distillation, and dynamic recompilation. We instantiate this paradigm as Skill Training. On FeatureBench, automatically discovered skills improve the passed rate by 8.1pp, and after progressive distillation the model still achieves a 27.7% passed rate without any external scaffold (distillation retention rate 85.2%, defined as post-distillation / with-skill passed rate), significantly outperforming standard SFT on the same data.
发表机构
- Alibaba Group(阿里巴巴集团)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。