发表机构
The University of Texas at Austin; Dell Pediatric Research Institute; College of Pharmacy; Chandra Family Department of Electrical and Computer Engineering; Department of Computer Science; College of Natural Sciences(德克萨斯大学奥斯汀分校; 戴尔儿科研究所; 药学院; 钱德拉家族电气与计算机工程系; 计算机科学系; 自然科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示高增益门控FFN行在文本与基因组基础模型中作为富集信号而非因果机制,其结构显著性可识别关键行,但因果组织因模型而异。
AI 中文摘要
少数异常高增益的参数能在Transformer语言模型中产生不成比例的影响,但类似结构是否在基因组基础模型中重现,以及结构几何是否决定功能重要性,仍属未知。我们分析了文本和基因组基础模型中门控前馈网络的高增益行,包括一个冻结的22模型因果普查。通过精确计算关联的双线性权重算子(不采用对角近似),我们检验了结构极端性是否是一种可迁移的机制。基于激活的候选行相对于随机和同层顶范数对照在功能上富集,但谱集中度和算子幅度均不能预测因果效应大小,且这些关联在端点同质的文本解码器子集内消失。在基因组和文本解码器各一个中,对36行进行层内扫描,将其解析为两种状态:低于检测器接受阈值时,比率不携带关于因果损伤的正向信息;而高于阈值时,比率能强排序行,但不以剂量-反应方式分级严重程度。同一扫描揭示了一个单候选每模型普查不可见的第二个单独灾难性行,以及共定位关键行之间的非加性损伤。案例研究显示了不同的因果组织:DNABERT-2中一个稳健的超加性成对交互,以及GENERator中一个尖锐位置局部化的依赖性,其中保留或恢复该行序列起始贡献可挽救几乎所有原生损失损伤。因此,高增益门控FFN行是一种反复出现的架构表型,其结构显著性充当富集信号,而非功能关键性的校准度量或因果组织的规范。富集是普遍的,但机制是模型特定的。
英文摘要
A small number of unusually high-gain parameters can exert disproportionate effects in transformer language models, but whether analogous structures recur in genomic foundation models and whether structural geometry determines functional importance remains unknown. We analyzed high-gain rows in gated feed-forward networks across text and genomic foundation models, including a frozen 22-model causal census. Computing an associated bilinear weight operator exactly, without a diagonal approximation, we tested whether structural extremeness is a transferable mechanism. Activation-derived candidates were functionally enriched relative to random and top-norm same-layer controls, yet neither spectral concentration nor operator magnitude predicted causal effect size, and these associations vanished within the endpoint-homogeneous text-decoder subset. A within-layer sweep of 36 rows in one genomic and one text decoder resolved this into two regimes: below the detector's acceptance threshold the ratio carried no positive information about causal damage, whereas above it the ratio ordered rows strongly but did not grade severity as a dose-response. The same sweep revealed a second individually catastrophic row invisible to a one-candidate-per-model census, and non-additive damage among co-located critical rows. Case studies showed divergent causal organizations: a robust super-additive pair interaction in DNABERT-2, and in GENERator a sharply position-localized dependence in which preserving or restoring the row's beginning-of-sequence contribution rescued essentially all native-loss damage. High-gain gated-FFN rows are therefore a recurrent architectural phenotype whose structural prominence acts as an enrichment signal, not a calibrated measure of functional criticality or a specification of causal organization. Enrichment is general, but the mechanism is model-specific.
Comments25 pages, 5 main figures, 2 supplementary figures