arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01835cs.AI

重写还是重加权?语言模型中的几何视角

Rewriting or Reweighting? A Geometric Account in Language Models

Juntong Wang, Shengkun Yang, Xiyuan Wang, Muhan Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出行为流形分析方法,通过ACT与NOC空间的几何视角,发现监督微调重写语言模型行为几何、奖励优化重加权该几何的机制差异,为理解后训练目标提供统一框架。

中文摘要 AI 辅助

后训练可大幅改变语言模型的行为,但整体行为率无法揭示训练是移除了现有机制、创建了新机制,还是改变了继承机制的使用方式。我们通过两种机制上不同的失败案例研究该问题:重复作为解码吸引子的病理现象,以及谄媚作为与偏好相关的对齐失败。我们引入行为流形分析,该方法通过选择与行为相关的稀疏坐标并将其提升到低维局部图中,来分离特定行为的几何结构。我们在两个互补空间中构建这些图:ACT捕获运行时激活状态,NOC量化模型通过共享行为相关子空间路由功能信息流的强度。在多个模型族中,生成的图高度压缩且在不同架构间部分可对齐。贡献空间图展现出更具架构鲁棒性的共享核心,而激活空间图则保留更强的特定模型族结构。通过受控后训练追踪这些图,我们发现一致的不对称性:监督微调(SFT)大幅改变继承的行为几何结构,而奖励优化改变行为的同时在很大程度上保留底层图。这种几何视角为理解两个目标间的机制差异提供了统一框架:SFT倾向于重写行为几何,而奖励优化主要对其进行重加权。代码可在该httpsURL获取。

英文摘要

Post-training can substantially alter language-model behavior, yet aggregate behavior rates do not reveal whether training removes an existing mechanism, creates a new one, or changes how an inherited mechanism is used. We study this question through two mechanistically distinct failures, repetition as a decoding-attractor pathology and sycophancy as a preference-related alignment failure. We introduce behavioral manifold analysis, which isolates behavior-specific geometry by selecting sparse behavior-associated coordinates and lifting them into low-dimensional local charts. We construct these charts in two complementary spaces. ACT captures runtime activation states, while NOC quantifies how strongly the model routes functional information flow through the shared behavior-associated subspace. Across multiple model families, the resulting charts are highly compressed and partially alignable across architectures. Contribution-space charts expose a more architecture-robust shared core, whereas activation-space charts retain stronger family-specific structure. Tracking these charts through controlled post-training reveals a consistent asymmetry. Supervised fine-tuning substantially alters the inherited behavioral geometry, whereas reward optimization changes behavior while largely preserving the underlying chart. This geometric perspective provides a unified framework for understanding the mechanistic distinction between the two objectives. SFT tends to rewrite behavioral geometry, whereas reward optimization primarily reweights it. Code is available at https://github.com/ronglingze/Manifold-Analysis

↑