发表机构
School of Computer Science and Technology, Xi’an Jiaotong University; Zhongguancun Academy, Beijing, China; National Engineering Research Center for Visual Information and Applications(西安交通大学计算机科学与技术学院; 北京中关村学院; 视觉信息与应用国家工程研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大模型后训练采用单一配方的问题,提出Self-Routing框架,通过样本行为状态路由至不同优化方式,在Qwen系列模型数学推理任务中性能优于基线,减少不必要更新。
AI 中文摘要
大型语言模型的后训练通常对所有样本采用单一训练配方,即便模型自身的滚动输出(rollouts)会呈现不同的样本级学习状态。我们提出Self-Routing,一种基于行为的后训练框架,利用滚动输出的正确性和置信度来决定每个样本的优化方式。根据样本的行为状态,它会被路由至GRPO、在线策略自蒸馏(on-policy self-distillation)、正则化或跳过,使训练无需外部教师、额外标注或额外采样即可自适应。在Qwen3和Qwen3.5主干的数学推理实验中,Self-Routing持续优于统一GRPO、统一OPSD、固定混合及更简单的路由基线。进一步分析显示,路由分布随训练变化,减少了对低信号或已稳定样本的不必要更新。
英文摘要
Post-training large language models usually applies a single training recipe to all samples, even though the model's own rollouts reveal different sample-level learning states. We propose Self-Routing, a behavior-conditioned post-training framework that uses rollout correctness and confidence to decide how each sample should be optimized. Depending on its behavior state, a sample is routed to GRPO, on-policy self-distillation, regularization, or skipping, allowing training to adapt without external teachers, extra annotations, or additional sampling. Experiments on mathematical reasoning across Qwen3 and Qwen3.5 backbones show that Self-Routing consistently improves over uniform GRPO, uniform OPSD, fixed mixtures, and simpler routing baselines. Further analyses show that the routing distribution changes over training and reduces unnecessary updates on low-signal or already stable samples.
Comments14 pages, 5 figures. Accepted at EMNLP 2026