使用奇偶瓶颈层扩展可解释的Transformer
Scaling Interpretable Transformers with Parity Bottleneck Layers
浏览论文内容
中文总结 AI 辅助
研究旨在解决训练可解释Transformer模型的难题,提出ParityTransformer架构,通过奇偶瓶颈层实现,实验表明其在稀疏探测任务中表现良好,在特征相关指标上更优,朝着训练内部表示可设计解释的模型迈出一步。
中文摘要 AI 辅助
语言模型被认为具有叠加现象,其残差流中表示的特征比维度更多。稀疏自动编码器旨在事后恢复这些特征,但训练具有可解释结构的模型仍然不切实际,因为每层的过完备瓶颈在内存和计算方面成本过高。为克服此问题,我们引入了ParityTransformer,这是一种GPT - 2规模的架构,其中间表示在设计上高效且宽/稀疏。每层用无参数代数字典替换学习到的过完备基,提供确定性的不相干性保证并消除内存需求。实验表明,ParityTransformer在稀疏探测任务上至少与事后稀疏自动编码器表现相当,在特征吸收、引导有效性和细粒度因果干预等方面表现更优,朝着训练内部表示可设计解释而非事后恢复的模型迈进。
英文摘要
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Sparse autoencoders (SAEs) are designed to recover such features post-hoc, but training models that are interpretable by construction has remained impractical, as a per-layer over-complete bottleneck is prohibitively expensive in both memory and compute. To overcome this issue, we introduce the ParityTransformer, a GPT-2-scale architecture whose intermediate representations are efficient and wide / sparse by design. At each layer, a Deep Parity Bottleneck (DPB) replaces a learned over-complete basis with a parameter-free algebraic dictionary, providing a deterministic incoherence guarantee and eliminating the memory requirements that have prevented per-layer interpretable bottlenecks at scale. A DPB is a hierarchically structured sparse bottleneck which efficiently enforces sparsity using a multi-level mixture-of-experts approach: a hardware-aware implementation that closes the cost gap between activation sparse and dense training to a manageable interpretability tax. Empirically, ParityTransformers perform at least as well as post-hoc SAEs on sparse probing tasks, while out-performing on measures of feature absorption, steering effectiveness, and fine-grained causal interventions. Because subsequent computation acts only on features that survive the sparse bottleneck, the ParityTransformer's features are native to the model's forwards pass by construction, addressing the question of whether SAEs probe features the model actually uses during computation. We see this as a step toward training models whose internal representations are interpretable by design rather than recovered post hoc.
发表机构
- Principles of Intelligence(智能原理)
机构由 AI 辅助整理,请以论文原文为准。