发表机构
University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究独立训练的语言模型在代码表示上的情况,通过2x2设计扩展概念电路提取方法,测量语法概念,发现任务决定概念获专用电路,模型决定电路位置及增长方式,还得出如Rust与Python电路差异等结论,为后续测试奠定基础。
AI 中文摘要
独立训练的语言模型是否会以相同的方式表示相同的事物?我们针对代码给出答案,将最近引入的概念电路提取方法扩展到2x2设计——Python和Rust与Qwen2.5-Coder-7B和DeepSeek-Coder-V1-6.7B交叉,并在所有四个单元中相同地测量完整的语法概念清单(58个Python,57个Rust):这是区分任务、语言和模型依赖因素的最小设计。答案分为三个部分。由任务决定哪些概念获得专用电路:模型在哪些概念接收电路上达成一致(Python的斯皮尔曼ρ = 0.638,Rust的为0.673,两者p < $10^{-7}$)。电路所在位置由模型决定:对于两种语言,Qwen在较晚频段(约L17 - 19)处理概念,DeepSeek在L6 - 7处理。电路跨层增长方式也由模型决定:Qwen赋予其原子概念早期峰值,而DeepSeek没有。因此,“电路是通用的吗?”没有单一答案:对于“是什么”是肯定的,对于“在哪里”和“如何”是否定的——通用性是表示内容的属性,而非计算组织的属性。所有这些结构都不是预先固定的。一致性可能落在独立和相同之间的任何位置;它落在ρ≈0.65处。在两个模型中,Rust结构比其Python对应结构接收多2 - 3倍的特定概念电路。两种模型在语言之间共享神经元(6/7和7/7配对结构),DeepSeek比Qwen多1.94倍——这一方向没有先前结果预测。并且Qwen将Rust类型和特征机制的九个关键字绑定到一个紧密的神经元簇中(杰卡德系数0.535对比空值0.112,p < 0.001),这是表面语法中不可见的语义维度。消融和线性探针证实电路是有功能的。所有断言都局限于这个2x2设计;每个模型的配置文件是否能预测第三个模型是接下来设计要测试的内容。
英文摘要
Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-circuit extraction method to a 2x2 design -- Python and Rust crossed with Qwen2.5-Coder-7B and DeepSeek-Coder-V1-6.7B -- and measuring a complete inventory of grammatical concepts (58 Python, 57 Rust) identically in all four cells: the smallest design that separates what depends on the task, the language, and the model. The answer splits into three parts. What earns dedicated circuitry is set by the task: the models agree on which concepts receive circuits (Spearman $ρ$ = 0.638 for Python, 0.673 for Rust, both p < $10^{-7}$). Where those circuits sit is set by the model: Qwen processes concepts in a late band (~L17-19), DeepSeek at L6-7, for both languages. How circuits grow across layers is also set by the model: Qwen gives its atomic concepts an early spike that DeepSeek does not. "Are circuits universal?" thus has no single answer: yes for What, no for Where and How -- universality is a property of representational content, not of computational organisation. None of this structure was fixed in advance. The agreement could have landed anywhere between independence and identity; it lands at $ρ\approx 0.65$. Rust constructs receive 2-3x more concept-specific circuitry than their Python equivalents, in both models. Both models share neurons between the languages (6/7 and 7/7 paired constructs), DeepSeek 1.94x more than Qwen -- a direction no prior result predicts. And Qwen binds nine keywords of Rust's type-and-trait machinery into one tight neuron cluster (Jaccard 0.535 vs null 0.112, p < 0.001), a semantic dimension invisible in surface syntax. Ablation and linear probes confirm the circuits are functional. All claims are scoped to this 2x2; whether the per-model profile predicts a third model is the designed next test.
Comments16 pages, 11 figures, 6 tables. Code: https://github.com/piotrwilam/Atlas2x2 ; dataset: https://huggingface.co/datasets/piotrwilam/Atlas2x2