重新思考权重绑定:用于语言模型稳定训练和更新的伪逆绑定
Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates
查看机构详情
- Monash University(墨尔本大学)
- Technical University of Munich(慕尼黑技术大学)
- Chongqing University(重庆大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出伪逆绑定方法,通过同步嵌入和解嵌入作为共享的潜在令牌记忆的耦合投影,提升语言模型训练稳定性与接口一致性,同时改善可解释性探查。
中文摘要 AI 辅助
权重绑定在紧凑语言模型中被广泛使用,以通过共享令牌表来减少参数。然而,仅参数共享并不能保证稳定的令牌接口:在训练过程中,将编码令牌转换为隐藏状态与将隐藏状态转换为logits的对应关系可能会漂移,从而恶化优化敏感性和依赖有意义词汇空间解码器的可解释性探查。我们提出伪逆绑定(PIT),通过将嵌入和解嵌入作为共享潜在令牌记忆的耦合投影,保证整个训练过程中的伪逆一致接口。PIT通过极化初始化从源检查点获得或thonormal共享内存,用于持续预训练,或通过随机或thonormal初始化用于从头预训练,并引入一个通过Cholesky因子参数化的学习对称正定隐藏空间转换参数。输出头在词汇投影之前应用此转换,而嵌入则使用稳定的三角求解应用反向转换,避免显式伪逆重新计算和词汇规模的辅助参数。除了提高训练稳定性外,PIT通过保持输入和输出令牌几何结构同步,为logit-lens风格和词汇空间可解释性探查提供了更清晰的基质。我们在设备模型上评估了PIT,参数范围从256M到1.3B。结果表明,PIT提高了持续预训练的稳定性,强制在不同设置下近似精确的令牌接口一致性,并在持续预训练后产生更可预测的轻量级适应,而从头预训练揭示了严格接口一致性与无约束优化之间的权衡。
英文摘要
Weight tying is widely used in compact language models to reduce parameters by sharing the token table between the input embedding and the output projection. However, parameter sharing alone does not guarantee a stable token interface: during training, the correspondence between encoding tokens into hidden states and decoding hidden states into logits can drift, worsening optimization sensitivity and weakening explainability probes that rely on a meaningful vocabulary-space decoder. We propose Pseudo-Inverse Tying (PIT), which synchronizes embedding and unembedding as coupled projections of a shared latent token memory, guaranteeing a pseudo-inverse-consistent interface throughout training. PIT maintains an orthonormal shared memory, obtained by polar initialization from a source checkpoint for continued pretraining or by random orthonormal initialization for from-scratch pretraining, and introduces a learned symmetric positive definite hidden-space transform parameterized via a Cholesky factor. The output head applies this transform to hidden states before the vocabulary projection, while the embedding applies the inverse transform to token vectors using stable triangular solves, avoiding explicit pseudo-inverse recomputation and vocabulary-sized auxiliary parameters. Beyond improving training stability, PIT provides a cleaner substrate for logit-lens-style and vocabulary-space explainability probes by keeping the input and output token geometries synchronized. We evaluate PIT on on-device models spanning 256M-1.3B parameters. The results show that PIT improves continued-pretraining stability, enforces near-exact token-interface consistency across settings, and yields more predictable lightweight adaptation after continued pretraining, while from-scratch pretraining reveals a trade-off between strict interface consistency and unconstrained optimization.