arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32913cs.LGcs.AI

Allspark:通过交替思维链实现从弱到强的迁移

Allspark: Weak to Strong Transfer via Alternating Chain of Thought

Kaizhao Liang, Junxiong Wang, Chen Liang, Zhendong Wang, Qiang Liu

AI总结:

Allspark 提出一种弱到强迁移框架,通过弱教师与冻结模型交替思维链训练,使强学生模型无需自身 rollout 即可获得推理能力提升,并在多个模型家族上验证了准确率提升。

AI中文摘要:

前沿模型的最新进展重新激发了人们对大规模强化学习(RL)的兴趣,但生成大模型 rollout 的成本使得即使测试 RL 配方也变得昂贵。我们探究了由小型弱模型学到的推理改进是否能使更大更强的模型受益,而无需在训练期间使用强模型的 rollout。我们提出了 Allspark,一个通过交替思维链实现从弱到强迁移的训练和推理框架。一个弱教师模型与同一模型的冻结副本一起训练;两者交替生成推理片段,冻结模型产生最终答案。在推理时,一个更强的学生模型取代冻结的训练伙伴,而两个模型均保持固定。由于它们通过文本进行通信,教师模型可以引导来自不同模型家族且具有不同分词器的学生模型。我们在两个规模上研究 Allspark:受控的 Qwen 实验(涵盖数学和推理任务),以及更大规模的 Inkling 实验(在 ARC-AGI-2 上进行)。Inkling 实验显示了在家族内和跨家族设置中的准确率提升,包括向 Kimi 和 Nemotron 的迁移,其收益因推理设置而异。这些发现激励了在多个强学生模型之间复用训练好的弱教师模型,并考察由此产生的准确率-计算量权衡。

英文摘要:

Recent progress in frontier models has renewed interest in large-scale reinforcement learning (RL), but the cost of generating large-model rollouts makes even testing RL recipes expensive. We ask whether reasoning improvements learned by a small, weak model can benefit a larger, stronger model without using the strong model's rollouts during training. We introduce Allspark, a training and inference framework for weak-to-strong transfer through alternating chains of thought. A weak teacher is trained alongside a frozen copy of the same model; the two alternate reasoning segments, and the frozen model produces the final answer. At inference time, a stronger student replaces the frozen training partner, while both models remain fixed. Because they communicate through text, the teacher can steer students from different model families and with different tokenizers. We study Allspark at two scales: controlled Qwen experiments across math and reasoning, and larger-scale Inkling experiments on ARC-AGI-2. The Inkling experiments show accuracy gains in within-family and cross-family settings, including transfer to Kimi and Nemotron, with benefits that vary across inference settings. These findings motivate reusing a trained weak teacher across strong students and examining the resulting accuracy--token tradeoff.

↑