Aha-Flow 蒸馏:流标记在大语言模型推理中的重要性
Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning
浏览论文内容
中文总结 AI 辅助
本文提出 Aha-Flow 蒸馏(AFD),通过识别流时刻和啊哈时刻,构建 Flow-CoT 并采用双模式自蒸馏,在 AIME25 和 HMMT25 上显著提升 Qwen3 模型的推理性能。
中文摘要 AI 辅助
我们识别出“流时刻”(Flow Moment),这是一种以持续的、确认过程的言语表达(如“我正在做”)为特征的推理模式,与以修订和回溯为导向的“啊哈时刻”(Aha Moment)形成对比。我们将它们对应的语言表达分别称为“流标记”(Flow Markers)和“啊哈标记”(Aha Markers)。基于这一观察,我们通过重写原始推理轨迹中的话语标记,同时保留其底层推理内容,构建了 Flow-CoT,并将其用作在线策略自蒸馏(OPSD)的辅助监督。我们进一步提出 \textbf{Aha-Flow 蒸馏(AFD)},这是 OPSD 的一种双模式扩展,将不同形式的特权信息与相应的推理指令配对。Aha 分支保留基于简洁解决方案的监督,而 Flow 分支在直接且自信的推理指令下引入重写的 Flow-CoT。在推理时,模型仅使用标准的反思指令,因此 Flow 风格的推理纯粹作为训练信号。在 AIME25 和 HMMT25 上的实验显示,在 Qwen3-8B 和 Qwen3-4B 上均取得一致改进:与复现的 OPSD 基线相比,AFD 在 Qwen3-8B 上将 Avg@12 从 60.8 提升至 61.3,在 Qwen3-4B 上从 57.5 提升至 58.6。受控消融进一步表明,在相同的 Flow-CoT/Aha-CoT 组成下,双模式训练将 Avg@12 从 59.5 提升至 60.1,表明收益不仅来自引入异构推理监督,还来自其在自蒸馏过程中的组织方式。代码可在 https://this https URL 获取。
英文摘要
We identify the Flow Moment, a reasoning pattern characterized by sustained, process-confirming verbalizations such as I'm doing, in contrast to the revision- and backtracking-oriented Aha Moment. We refer to their corresponding linguistic expressions as Flow Markers and Aha Markers, respectively. Based on this observation, we construct Flow-CoT by rewriting the discourse markers of original reasoning traces while preserving their underlying reasoning content, and use it as auxiliary supervision for on-policy self-distillation (OPSD). We further propose \textbf{Aha-Flow Distillation (AFD)}, a dual-mode extension of OPSD that pairs different forms of privileged information with corresponding reasoning instructions. The Aha branch retains concise solution-based supervision, while the Flow branch introduces rewritten Flow-CoT under a direct and confident reasoning instruction. At inference time, the model uses only the standard reflective instruction, so Flow-style reasoning serves purely as a training signal. Experiments on AIME25 and HMMT25 show consistent improvements across Qwen3-8B and Qwen3-4B: AFD improves Avg@12 from 60.8 to 61.3 on Qwen3-8B and from 57.5 to 58.6 on Qwen3-4B over our reproduced OPSD baselines. Controlled ablations further show that, with the same Flow-CoT/Aha-CoT composition, dual-mode training improves Avg@12 from 59.5 to 60.1, indicating that the benefit comes not only from introducing heterogeneous reasoning supervision, but also from how it is organized during self-distillation. The code is available at https://github.com/Wang-Xiaodong1899/Aha-Flow-Distillation.
发表机构
- Peking University(北京大学)
- Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。