稀疏自编码器相图中的主导弥散相
A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram
浏览论文内容
中文总结 AI 辅助
本研究通过MAIS-O43协议实验发现,稀疏自编码器训练收敛于弥散相而非字典恢复或特征合并,表明训练模型的相图与目标最优解存在根本差异。
中文摘要 AI 辅助
稀疏自编码器(SAEs)日益被用于从神经网络激活中恢复可解释特征,然而系统性的特征共现可能导致不同特征被吸收或合并。MAIS-O43开放问题提出了一项受控实验,以刻画当嵌套比例 $\gamma$、稀疏惩罚 $\lambda$ 和字典大小 $M$ 变化时,真实合成字典的恢复何时让位于特征合并。我们实现了指定协议,并在165个网格单元中的10个上评估了200次独立初始化的拟合。我们观察到零次全字典恢复和零次合并。相反,每次运行都收敛到一个可复现的弥散相:重建几乎完美,但学习到的原子通常远离真实特征(中位最佳余弦为0.5-0.7,而恢复标准为0.95),且学习到的编码比真实编码密集一个数量级。在稳健性检查以及使用标准小批量Adam(额外3300次拟合)的完整165单元网格中,该行为持续存在。由于精确稀疏编码目标的全局最优值在双特征情况下已知会合并嵌套特征,这些结果表明,训练后的SAE不必达到相应的最小值,且训练模型的相图可能在根本上不同于目标最小化器的相图。
英文摘要
Sparse autoencoders (SAEs) are increasingly used to recover interpretable features from neural-network activations, yet systematic feature co-occurrence can cause distinct features to be absorbed or merged. The MAIS-O43 open problem proposes a controlled experiment to characterize when recovery of a true synthetic dictionary gives way to feature merging as the nesting fraction $γ$, sparsity penalty $λ$, and dictionary size $M$ vary. We implement the specified protocol and evaluate 200 independently initialized fits across ten of the 165 grid cells. We observe zero full-dictionary recoveries and zero merges. Instead, every run converges to a reproducible diffuse phase: reconstruction is nearly perfect, but learned atoms typically remain far from the true features (median best cosine 0.5-0.7 against a 0.95 recovery criterion) and learned codes are an order of magnitude denser than the ground truth. This behavior persists under robustness checks and across the full 165-cell grid using standard minibatch Adam (3,300 additional fits). Since the global optimum of the exact sparse-coding objective is known to merge nested features in the two-feature case, these results suggest that trained SAEs need not reach the corresponding minima, and that the phase diagram of trained models may differ fundamentally from that of objective minimizers.