arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09058cs.LG

学习跨模型激活对齐与显式多对多层映射

Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps

  • Technion(以色列理工学院)
  • NVIDIA Research(英伟达研究院)

机构由 AI 辅助整理,请以论文原文为准。

Alina Sudakov, Guy Bar-Shalom, Fabrizio Frasca, Haggai Maron

AI总结:

本研究提出MATCHA方法,通过联合学习层映射与特征映射,实现跨模型激活对齐,在42对模型上提升重建与检索指标,并支持干预迁移。

AI中文摘要:

大型语言模型(LLMs)以快速步伐发布,引发了一个自然的问题:两个独立训练的模型之间如何关联,既包括哪些层对应,也包括特征如何在它们之间转换。我们通过学习激活对齐(activation alignment)来研究这一点,即从源模型的逐层激活到目标模型的映射。我们的方法MATCHA将此映射分解为层映射(layer map),其输出是一个显式的目标-源矩阵,可以提取和检查,以及一个层共享的特征映射(feature map),位于隐藏空间之间。大多数先前的工作预先固定层对应关系,将大致相同相对深度的层配对;相比之下,我们从提示中联合学习两个因素。在跨越三个不同家族的七个模型的42对组合中,MATCHA更忠实地重建了目标的激活,并在检索式指标上大幅提升,相对于先前的方法。恢复的映射在深度上大致单调,但与大多数先前方法不同,它们始终是多对多的:每个目标层都利用一个源层带。我们的对齐还支持激活空间干预的迁移,使得为一个模型开发的转向向量(steering vectors)和探针(probes)能够迁移到另一个模型。

英文摘要:

LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them? We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's. Our method, MATCHA, factors this map into a layer map, whose output is an explicit target-by-source matrix that can be extracted and inspected, and a layer-shared feature map between hidden spaces. Most of prior work fixes the layer correspondence in advance, pairing layers at roughly the same relative depth; in contrast, we learn both factors jointly from prompts. Across 42 pairs of seven models spanning three different families, MATCHA reconstructs the target's activations more faithfully and improves retrieval-based metrics substantially, w.r.t. previous approaches. The recovered maps are broadly monotone in depth but, in contrast with most previous approaches, are consistently many-to-many: each target layer draws on a band of source layers. Our alignments also enable transfer of activation-space interventions, allowing steering vectors and probes developed for one model to transfer to another.

↑