超越压缩:诊断后训练如何改变数学推理
Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning
- University of Luxembourg(卢森堡大学)
- Seafill Open-Source Community(Seafill 开源社区)
- Université Paris-Saclay(巴黎-萨克雷大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过跨表面pass@K等诊断方法比较三种后训练路径,发现后训练在简单问题上主要压缩采样成本,在难题上扩大上限,且压缩并非后训练的普遍机制。
AI中文摘要:
后训练是现代大型语言模型(LLMs)数学推理的核心环节,但仅凭端点的pass@1指标不足以识别发生了什么变化。性能提升可能源于新可达的解决方案、对潜在解决方案的更廉价采样、表面鲁棒性或记忆效应。我们在统一的诊断读数下比较了三条后训练路径:我们充分训练的离策略蒸馏轨迹、发布的Qwen3离策略加在策略蒸馏端点,以及使用组相对策略优化(GRPO)训练的已发布DeepSeek-Math端点。我们的探针使用跨表面的pass@K,涵盖逐字提示、释义、数值同构和翻译,并辅以一致性、分布形状和验证的监督微调(SFT)成员分析。我们发现两种机制。在较简单的AMC问题上,大K上限接近饱和,因此后训练主要压缩采样成本。在较难的AIME问题上,后训练相对于基础模型扩大了大K上限:充分的离策略蒸馏已能提高此上限,Qwen3发布端点进一步提高了它,而DeepSeek-Math GRPO在大K下并未优于充分的离策略蒸馏。英语主导的蒸馏改善了非英语推理,但保留了语言层级差距。一项受控过拟合审计发现当前SFT成员探针的敏感性有限。压缩是后训练的一种机制,而非普遍解释。
英文摘要:
Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.