arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型中的跨架构引导迁移:一项系统实证研究

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

Ayushi Agarwal

arXiv 2608.05164首次发表:更新:

AI 中文总结

本研究通过系统实证发现,在1.7B参数规模以上,独立训练的大语言模型间存在可利用的跨架构引导迁移,验证了柏拉图式表征假说的功能层面补充,强调了机制可解释性中的规模阈值。

AI 中文摘要

独立训练的大型语言模型尽管架构存在差异,仍可能形成语义概念的共享内部表征,但这种几何相似性是否对跨模型行为控制具有功能影响尚未得到验证。本文首次对跨模型引导迁移进行系统评估,结果表明共享的大语言模型(LLM)几何结构具有功能可利用性,但存在条件限制:当具备足够的表征能力时,一个模型的概念方向可引导另一个独立训练的模型。本研究选取5个开放权重模型,涵盖3种参数规模(0.8B至8B)和2种架构谱系,为每个模型在15个语义域上训练一个稀疏自编码器(Sparse Autoencoder),并在全部20个有向模型对上测试对齐情况。我们观察到在1.7B参数附近存在显著的不连续性:在≥1.7B规模时,47%至49%的跨模型特征对通过验证(皮尔逊相关系数r≥0.60,普罗克拉斯提斯余弦值为0.895至0.956),而在0.8B以下时对齐情况急剧下降。跨模型引导向量(B3-TI)在15个监督概念上的胜率达71.0%,而同模型原生向量的胜率为68.0%;单个通用向量在5个模型中的4个上取得67.3%的胜率,无需任何针对单个模型的监督。迁移效果在低于1.7B的模型以及存在生成不稳定性的模型中会下降,这证实了功能可利用性需要足够的表征能力。我们的研究强调了机制可解释性中规模阈值的重要性:在7B规模下验证的工具若不重新验证,可能无法迁移到更小的模型。本文为柏拉图式表征假说提供了首个功能层面的补充:在确定的规模条件下,独立训练的LLM之间的几何收敛支持无需微调的跨模型行为控制。

英文摘要

Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences -- but whether this geometric similarity has functional consequences for cross-model behavioural control remains untested. We present the first systematic evaluation of cross-model steering transfer and show that shared LLM geometry is functionally exploitable, conditionally: concept directions from one model can steer a different independently trained model when sufficient representational capacity exists. We study five open-weight models spanning three parameter scales (0.8B--8B) and two architectural lineages, training one Sparse Autoencoder per model across 15 semantic domains and testing alignment across all 20 directed model pairs. We observe a suggestive discontinuity near 1.7B parameters: at >= 1.7B scale, 47--49% of cross-model feature pairs validate (Pearson r >= 0.60, Procrustes cosines 0.895--0.956), while alignment degrades sharply below 0.8B. Cross-model steering vectors (B3-TI) achieve a 71.0% win rate across 15 supervised concepts versus 68.0% for same-model native vectors; a single universal vector achieves 67.3% in 4 of 5 models without any per-model supervision. Transfer degrades for models below 1.7B and for one model with generation instability, confirming that functional exploitability requires sufficient representational capacity. Our findings underscore the importance of scale thresholds in mechanistic interpretability: tools validated at 7B scale may not transfer to smaller models without revalidation. We provide the first functional complement to the Platonic Representation Hypothesis -- geometric convergence across independently trained LLMs supports cross-model behavioural control without fine-tuning, under the identified scale conditions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑