超越固定方向:大语言模型中推理与记忆的自适应表示分析
Beyond Fixed Directions: Adaptive Representation Analysis of Reasoning and Memorization in LLMs
浏览论文内容
中文总结 AI 辅助
该研究以Qwen3-0.6B为对象,通过实验验证了推理与记忆可通过单一方向分离,但经GRPO后该方向几何结构显著重组,不过任务组的AUROC仍保持1.00,挑战了固定方向稳定性的假设。
中文摘要 AI 辅助
近期研究提出,语言模型中的推理与记忆可通过单一表示方向来表征,包括在强化学习过程中保持该方向固定的方法。我们检验了这一观点背后的两个假设:其一,面向推理的任务组与事实回忆任务组是否近似可通过单一方向分离;其二,在使用GRPO后,由此产生的几何结构是否仍保持稳定。我们使用Qwen3-0.6B和一个受控的400样本数据集进行研究,发现一维投影在研究的任务组上的AUROC可与完整1024维线性探针相匹配,达到1.00。然而,在GRPO后,对应的方向发生了显著重组:平均方向余弦平均值为0.453,探针方向余弦为0.445,而最终层的直接表示漂移达到0.511。尽管如此,探针的AUROC仍保持为1.00。因此,证据支持研究的任务组具有单一方向可解码性,但对固定方向的稳定性提出了挑战:信息得以保留,但其几何实现发生了变化。
英文摘要
Recent work has proposed that reasoning and memorization in language models can be characterized by a single representation direction, including methods that keep this direction fixed during reinforcement learning. We test two assumptions behind this view. First, are reasoning-oriented and factual-recall task groups approximately single-direction separable? Second, does the resulting geometry remain stable after GRPO? Using Qwen3-0.6B and a controlled 400-example dataset, we find that a one-dimensional projection can match a full 1024-dimensional linear probe with AUROC = 1.00 on the studied task groups. However, after GRPO, the corresponding direction is substantially reorganized: mean-direction cosine averages 0.453, probe-direction cosine 0.445, while direct representation drift reaches 0.511 at the final layer. Probe AUROC nevertheless remains 1.00. The evidence therefore supports single-direction decodability for the studied task groups but challenges fixed-direction stability: the information persists while its geometric realization changes.