arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DV-Lens:揭示语言模型参数的功能组织

DV-Lens: Revealing the Functional Organization of Language Model Parameters

Chenhang Cui, Jian Yu, Shuyi Miao, Xiaohao Liu, Rui Huang, Fei Shen, An Zhang, Tat-Seng Chua

arXiv 2610.04489首次发表:更新:

发表机构

National University of Singapore; Nanjing University of Science and Technology; Beihang University; University of Hong Kong; University of Science and Technology of China(新加坡国立大学; 南京理工大学; 北京航空航天大学; 香港大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DV-Lens参数级可解释框架,通过下游词汇响应连接参数与输出,并引入DV-Complexity度量,在48个模型上验证了其与模型能力的强相关性。

AI 中文摘要

理解参数功能有助于阐明大型语言模型(LLM)的内部机制。然而,如何将不同模块的参数与可验证的输出效果联系起来,并进一步刻画其功能组织与模型能力之间的关系,仍有待探索。为此,我们引入了下游词汇透镜(DV-Lens),一种参数级可解释性框架,将原生参数方向与其下游词汇响应联系起来。具体而言,我们首先针对注意力查询、键、值、输出(Q/K/V/O)投影和前馈网络(FFN),在参考提示集上估计模块特定的下游雅可比矩阵。其次,我们利用这些映射将原生参数列投影到最终词汇空间,获得有符号读数,以刻画其平均局部输出响应。第三,我们根据词汇读数对参数列进行分组,并引入下游词汇复杂度(DV-Complexity),利用原始权重的归一化重建残差量化组内结构变异。在参数层面,随机对照和有限差分测试表明,DV-Lens读数捕获了非随机的词汇结构,并在来自九个模型的720个案例中,以98.0%的坐标方向一致性预测局部逻辑值变化。这些读数进一步指导了21个模型上的参数消融、引导和交换,在受控条件下按预测方向改变目标标记概率。在模型层面,DV-Complexity的联合参数得分在48个语言模型上与基于基准的能力排名实现了0.904的斯皮尔曼相关性。这些结果共同为DV-Lens的解释提供了基于干预的证据,并揭示了DV-Complexity与模型能力之间的关联。

英文摘要

Understanding parameter functions helps elucidate the internal mechanisms of large language models (LLMs). However, how to connect parameters from different modules to verifiable output effects and further characterize the relationship between their functional organization and model capability remains to be explored. To this end, we introduce the downstream vocabulary lens (DV-Lens), a parameter-level interpretability framework that links native parameter directions to their downstream vocabulary responses. Specifically, we first estimate module-specific downstream Jacobians over a reference prompt set for attention query, key, value, and output (Q/K/V/O) projections and feed-forward networks (FFNs). Second, we use these mappings to project native parameter columns into the final vocabulary space, obtaining signed readouts that characterize their average local output responses. Third, we group parameter columns by their vocabulary readouts and introduce downstream vocabulary complexity (DV-Complexity), which quantifies within-group structural variation using normalized reconstruction residuals of the original weights. At the parameter level, randomized controls and finite-difference tests show that DV-Lens readouts capture non-random vocabulary structure and predict local logit changes with 98.0% coordinate-orientation agreement across 720 cases from nine models. These readouts further guide parameter ablation, steering, and swapping across 21 models, shifting target-token probabilities in the predicted directions under controlled conditions. At the model level, the joint-parameter score of DV-Complexity achieves a Spearman correlation of 0.904 with benchmark-based capability rankings across 48 language models. Together, these results provide intervention-based evidence for DV-Lens interpretations and reveal an association between DV-Complexity and model capability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑