在线策略注意力线性化
On-Policy Attention Linearization
AI总结:
针对混合注意力模型离线蒸馏在长上下文任务中性能崩溃的问题,提出在线策略注意力线性化(OPAL),让学生模型自采样长轨迹并接受教师密集监督,在仅用30亿token下恢复检索100%及推理大部分性能。
AI中文摘要:
混合Transformer架构用线性注意力替代大部分软注意力层,以极低的显存成本提供Transformer级别的质量。越来越多的研究并非从头预训练这类模型,而是从已训练好的全注意力Transformer中进行蒸馏。然而,这些蒸馏模型在长上下文检索和推理任务上经常崩溃,尤其是在思考模式下,而混合架构的效率优势在此时最为关键。由于线性注意力层必须将上下文压缩为固定大小的状态,其误差会随长序列累积。由于离线策略蒸馏从未教导学生模型从这种漂移中恢复,需要更长序列长度的任务变得尤为困难。我们提出了在线策略注意力线性化(OPAL),其中混合注意力学生模型自行采样其长上下文轨迹,并从冻结的全注意力教师模型获得密集监督。将OPAL应用于Qwen3-4B和MiMo-7B-RL-0530,我们恢复了全注意力模型在常识推理上87%–94%的性能,在针在干草堆(NIAH)检索上恢复100%,在数学推理上恢复83%–93%,仅使用30亿训练token。我们在没有监督微调(SFT)或可验证奖励的强化学习(RLVR)的情况下取得了这些结果。与之前最强线性化方法相比(该方法恢复了教师检索性能的68%和数学推理平均准确率的21.6个百分点),OPAL完全恢复了检索,并在数学推理上达到67.6%–72.2%。
英文摘要:
Hybrid transformer architectures that replace most softmax attention layers with linear attention offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained full-attention transformers. However, these distilled models often collapse on long-context retrieval and reasoning tasks, particularly when operating in thinking mode, where the efficiency gains of hybrid architectures matter most. Since linear attention layers must compress context into a fixed-size state, their errors compound over long sequences. As off-policy distillation never teaches the student model to recover from this drift, tasks that necessitate longer sequence lengths become especially challenging. We introduce On-Policy Attention Linearization (OPAL) in which the hybrid attention student samples its own long-context trajectories and receives dense supervision from the frozen full-attention teacher. Applying OPAL to Qwen3-4B and MiMo-7B-RL-0530, we recover $87$--$94\%$ of full-attention performance on commonsense reasoning, $100\%$ on needle-in-a-haystack (NIAH) retrieval, and $83$--$93\%$ on mathematical reasoning with only 3B training tokens. We achieve these results without supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR). Compared with the strongest prior linearization method, which recovers $68\%$ of its teacher's retrieval performance and $21.6\%$ absolute average mathematical reasoning accuracy, OPAL fully recovers retrieval and achieves $67.6$--$72.2\%$ on math reasoning.