arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16347cs.CLcs.LG

大语言模型间激活状态的架构依赖型因果迁移

Architecture-Dependent Causal Transfer of Activation States Across Large Language Models

Fernando Cardenas Piepereit

首次发表
浏览论文内容

中文总结 AI 辅助

本文探究大语言模型间激活状态的因果迁移,发现该迁移依赖模型架构,仅部分仅解码器模型对可实现具统计显著性的因果效应,且迁移的是表征载体而非意义。

中文摘要 AI 辅助

AI系统间的直接通信依赖自然语言作为中间层,会产生编码/解码开销、token成本和延迟。本文探究是否可通过学习得到的投影,在不同大语言模型(LLM)架构间因果迁移内部激活状态,从三个层面评估:表征相似性、从投影状态的跨模型检索,以及生成过程中通过激活注入的端到端因果迁移。使用四个架构各异的开放权重模型(Qwen2-0.5B、Phi-3-mini、Mistral-7B、FLAN-T5-base),研究发现训练后模型的表征对齐程度超过随机初始化的空基线,且可通过基于秩的指标(互k近邻对齐)最佳捕捉,该指标比中心核对齐(CKA)或Procrustes分析对激活幅度异常值更具鲁棒性。学习得到的投影网络从保留集中检索目标模型表征的表现远高于随机水平,三个仅解码器模型对的top-1准确率为45%-50%(随机水平为5%),但基于编码器的FLAN-T5则处于随机水平。生成过程中向目标模型注入投影激活,仅在三个仅解码器对中的一对(Qwen2-0.5B到Phi-3-mini)对基于检索的输出相似性产生统计显著的预先注册因果效应(23.3%,阴性对照下为0.0%,p=0.047,经FDR校正);尽管Mistral-7B相关的两对在隐状态层面具有可比的表征对齐,却未出现此类效应。本文将结果解读为表征载体的因果迁移而非意义的迁移,并得出结论:当前实现的LLM间端到端激活状态迁移是架构依赖的,而非通用的。

英文摘要

Direct communication between AI systems relies on natural language as an intermediate layer, incurring encoding/decoding overhead, token cost, and latency. We ask whether internal activation states can instead be transferred causally between different large language model (LLM) architectures via a learned projection, evaluated at three levels: representational similarity, cross-model retrieval from projected states, and end-to-end causal transfer via activation injection during generation. Using four architecturally diverse open-weight models (Qwen2-0.5B, Phi-3-mini, Mistral-7B, FLAN-T5-base), we find that representational alignment in trained models exceeds a random-initialization null baseline and is best captured by a rank-based metric (mutual k-nearest-neighbour alignment), more robust to activation-magnitude outliers than centered kernel alignment (CKA) or Procrustes analysis. A learned projection network retrieves the correct target-model representation from a held-out set well above chance for the three causal decoder-only model pairs (45-50% top-1 accuracy vs. 5% chance) but at chance level for the encoder-based FLAN-T5. Injecting projected activations into a target model during generation produces a statistically significant, pre-registered causal effect on retrieval-based output similarity for only one of the three decoder-only pairs (Qwen2-0.5B to Phi-3-mini: 23.3% vs. 0.0% under negative control, p=0.047, FDR-corrected); the two pairs targeting Mistral-7B show no such effect despite comparable representational alignment at the hidden-state level. We interpret these results as evidence for causal transfer of the representational vehicle, not of meaning, and conclude that end-to-end activation-state transfer between LLMs, as currently implemented, is architecture-dependent rather than universal.

补充信息

↑