发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究稀疏自编码器中哪些读出能在稀疏压缩中幸存的问题,核心方法是将最优岭预测算子间的矩阵值失真与重建损失结合训练解码器保留SAEs,主要贡献为得出重建质量不能决定读出幸存情况等结论。
AI 中文摘要
稀疏自编码器(SAEs)将模型激活压缩为稀疏码,但相同的重建误差和稀疏性可能保留不同的线性可解码信号。我们将这种模糊性形式化为最优岭预测算子之间的矩阵值失真,并通过将此失真与重建损失相结合来训练解码器保留SAEs。在秩松弛中,各向同性任务先验使每个模式的遗漏成本饱和而不改变主成分分析的排序,而结构化先验可改变保留的模式。一个受控稀疏实验表明声明的先验可保护留出的组合免受其任务子空间的影响。在GPT - 2小块8上,DPSAE在三次配对运行中使留出的解码器失真降低10.6 - 11.4%,同时匹配重建归一化均方误差。这些结果表明重建质量不能决定哪些重新拟合的线性读出能在稀疏压缩中幸存,且读出保留不同于学习更清晰的基准概念或保留每个冻结模型行为。
英文摘要
Sparse autoencoders (SAEs) compress model activations into sparse codes, but equal reconstruction error and sparsity can preserve different linearly decodable signals. We formalize this ambiguity as a matrix-valued distortion between optimal ridge-prediction operators and train decoder-preserving SAEs by combining this distortion with reconstruction loss. In a rank relaxation, an isotropic task prior saturates per-mode omission costs without changing PCA's ordering, whereas a structured prior can change which modes are retained. A controlled sparse experiment shows that a declared prior protects held-out combinations from its task subspace. On GPT-2 small block 8, DPSAE reduces held-out decoder distortion by 10.6--11.4% across three paired runs while matching reconstruction NMSE. The same checkpoints pass an average natural-text output-KL noninferiority test, but one matched Pythia pair shows no improvement in probes restricted to a few sparse features. These results show that reconstruction quality does not determine which refitted linear readouts survive sparse compression, and that readout preservation is distinct from learning cleaner benchmark concepts or preserving every frozen-model behavior.