AI 中文总结
本研究针对有限同构符号域,通过预先指定的构造-验证测试,在 Qwen2.5-7B-Instruct 模型的特定层区间,复现了某一候选路径的操作级因果迁移效应,但未确立更广泛的泛化性。
AI 中文摘要
行为准确率、线性可解码性及成功的激活干预本身并不能表明模型从一个符号域到另一个符号域具备操作级结构。我们在有限同构状态空间中提出一个更狭窄的问题:若针对每个源输入分别估计两个操作间的隐状态差异,将该差异添加到映射后的目标输入中,是否会使模型向对应的目标答案移动?该设计将这种输入特定干预与错误操作、规范匹配的随机操作及无操作对照进行比较,并将候选构造与独立隔离的验证拆分分开。在冻结的 Qwen2.5-7B-Instruct 模型的第20至21层,一个在验证访问前预先指定并冻结的候选路径-域-操作组合(transparent | integer_mod16--letters16 | successor->predecessor)通过了两次 PyVene 拆分;其验证交集-并集 p 值为 0.000198,36个家族的 Holm 校正 p 值为 0.006943。后续的 NNsight 0.7.0 实验在验证访问前预先指定并冻结,仅测试该选定的提示路径,未重新选择候选或层。该实验在数值上复现了全部12个验证效应估计值、置信区间及精确符号翻转 p 值;其36个家族的 Holm 校正 p 值为 0.007141。因此,该结果仅局限于一个提示路径和一个候选,在同一模型版本和同一层区间的两次干预实现中得到复现,并未确立跨模型泛化、全家族后端独立性、域通用迁移或代数不变性。
英文摘要
Behavioral accuracy, linear decodability, and successful activation interventions do not by themselves show that a model carries an operation-level structure from one symbolic domain to another. We ask a narrower question in finite isomorphic state spaces: if the hidden-state difference between two operations is estimated separately for each source input, does adding that difference to a mapped recipient input move the model toward the corresponding recipient answer? The design compares this input-specific intervention with wrong-operation, norm-matched random, and no-op controls, and separates candidate construction from an independently isolated confirmation split. On a frozen Qwen2.5-7B-Instruct model at layers 20--21, one route--domain--operation candidate from a family pre-specified and frozen before confirmation access, transparent | integer_mod16--letters16 | successor->predecessor, passed both PyVene splits; its confirmation intersection--union p-value was 0.000198 and its 36-family Holm-adjusted p-value was 0.006943. A subsequent NNsight 0.7.0 experiment, pre-specified and frozen before its confirmation access, tested only this selected prompt route, without candidate or layer reselection. It reproduced all 12 confirmation effect estimates, confidence intervals, and exact sign-flip p-values numerically; its 36-family Holm-adjusted p-value was 0.007141. The result is therefore limited to one prompt route and one candidate, replicated across two intervention implementations on one model revision and one layer interval. It does not establish cross-model generalization, full-family backend independence, domain-general transfer, or algebraic invariance.
Comments20 pages, 5 figures, and 4 ancillary CSV files