arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NinaXander:通过共享潜在空间跨架构家族组合冻结语言模型的可行性与局限

NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space

Takanori Kotama, Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri

arXiv 2609.38261首次发表:更新:

发表机构

Nagoya University(名古屋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出NinaXander,通过共享潜在适配器连接不同架构家族的冻结语言模型,验证跨家族组合的可行性,但发现组合模型在准确率和语言建模上不及父模型,且不共享通用语义空间。

AI 中文摘要

在本文中,我们提出了NinaXander,一系列组合语言模型,通过使用一个训练好的共享潜在适配器,将来自不同架构家族的冻结语言模型的层连接起来。组合模型运行一个模型的前几层,用适配器将中间表示转换一次,然后运行另一个模型的其余层。一旦适配器训练完成,就可以在不重新训练的情况下获得多个在不同层连接的组合模型。本研究使用循环的RWKV-4-Raven-7B和基于Transformer的Tulu-Pythia-6.9b(简称为RWKV和Pythia),考察了不同家族的冻结模型是否可以在事后重新组合。组合模型回答了多项选择题,我们检查的生成结果产生了语法良好的文本。将Pythia的前5层与RWKV的其余27层相结合的配置,将Transformer键值(KV)缓存减少了84.4%,准确率与单独使用RWKV相比无显著差异。然而,在多项选择题准确率方面,没有组合模型能达到父模型Pythia的水平,语言建模性能在WikiText(训练领域之外的维基百科文章语料库)上急剧下降。在一个有利的情况下,即共享分词器、相同深度和相同隐藏宽度时,也获得了中间表示的对应关系,但这并不表明模型共享一个通用的语义空间。

英文摘要

In this paper we propose NinaXander, a series of composed language models obtained by connecting layers of frozen language models from different architecture families with a single trained shared-latent adapter. A composed model runs the first layers of one model, converts the resulting intermediate representation once with the adapter, and then runs the remaining layers of the other model. Once the adapter is trained, several composed models that connect at different layers are obtained without retraining. Using the recurrent RWKV-4-Raven-7B and the Transformer-based Tulu-Pythia-6.9b, abbreviated as RWKV and Pythia, this study examines whether frozen models from different families can be recombined post hoc. The composed models answered multiple-choice questions, and those whose generations we examined produced syntactically well-formed text. The configuration that combines the first 5 layers of Pythia with the remaining 27 layers of RWKV reduced the Transformer key-value (KV) cache by 84.4% with accuracy not significantly different from that of RWKV alone. In multiple-choice accuracy, however, no composed model matched the parent model Pythia, and language-modeling performance decreased sharply on WikiText, a corpus of Wikipedia articles outside the training domain. The correspondence between intermediate representations was also obtained in one favorable case, with a shared tokenizer, the same depth, and the same hidden width, and does not show that the models share a general semantic space.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑