发表机构
Columbia University Irving Medical Center(哥伦比亚大学伊文思医疗中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对基因组语言模型内部表示不透明问题,引入结合稀疏字典学习与因果干预的框架,训练自动编码器提取转录因子结合特征,开发去混淆协议并因果验证,为基因组深度学习可解释性提供计算标准。
AI 中文摘要
基因组语言模型在调控基因组任务中表现出色,但模型内部表示仍不透明,缺乏验证模型中表观“概念”真实性的原则性程序。我们引入一个结合稀疏字典学习和因果干预的框架,在两个不同架构模型的隐藏激活上训练top-$k$稀疏自动编码器,恢复数千个映射到转录因子序列基序的单语义特征。针对位置权重矩阵对这些特征的朴素验证受GC组成和重复元素严重混淆,我们开发了消除这些混淆的协议。通过在模型前向传播中消融单个字典方向并测量模型预测分布的诱导变化,确定特定特征因果地用于表示细胞类型特异性转录因子结合。在三个转录因子和两种架构上,因果验证的结合特征可重复出现,而两类阴性对照无信号。该框架纯计算且仅使用公共数据,为基因组深度学习中的可解释性声明提供了可重复使用的标准。
英文摘要
Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition. We introduce a framework that combines sparse dictionary learning with causal intervention to extract, validate, and causally test interpretable features in genomic foundation models. Training top-$k$ sparse autoencoders on the hidden activations of two architecturally distinct models, Nucleotide Transformer ($6$-mer tokenization) and DNABERT-2 (byte-pair encoding), we recover thousands of monosemantic features that map to transcription-factor (TF) sequence motifs. We show that the naive validation of such features against position weight matrices is severely confounded by GC composition and repetitive elements, producing hundreds of spurious ``TF features'', and we develop a composition-matched, binding-resolved protocol that removes these confounds. Critically, we move beyond correlation: by ablating individual dictionary directions during the model's forward pass and measuring the induced shift in the model's own predictive distribution, we establish that specific features are \emph{causally} used to represent cell-type-specific TF binding, not merely motif presence. Across three transcription factors (CTCF, GATA1, REST) and both architectures, causally validated binding features emerge reproducibly ($7$--$14$ of $15$ tested features per condition), while two classes of negative control, scrambled binding labels and randomly selected features, yield no detectable signal. The framework is purely computational, uses only public data, and provides a reusable standard for interpretability claims in genomic deep learning.