当树不够用时:基于自适应图稀疏自编码器学习混合拓扑特征图
When Trees Are Not Enough: Learning Mixed-Topology Feature Graphs with Adaptive Graph Sparse Autoencoders
浏览论文内容
中文总结 AI 辅助
提出自适应图稀疏自编码器(AG-SAE),通过竞争完整父集学习混合拓扑特征图,以结构损失引导训练并闭环优化字典,实现比树结构更可靠的特征组织与更强因果干预。
中文摘要 AI 辅助
稀疏自编码器(SAE)能够揭示大型语言模型激活中的可解释特征,然而现有的结构化SAE强制采用单父树或森林结构,而事后图虽允许多父节点,但既不引导特征学习,也无法确保可靠的关联恢复。我们提出了自适应图稀疏自编码器(AG-SAE),这是一种结构引导的训练范式,将每个特征的完整父节点集合视为一个原子结构假设,并让证据选择零个、一个或多个父节点。通过将完整父节点集合与空集、子集及替代解释进行竞争,AG-SAE能够识别出联合必要的多父节点关联,同时拒绝冗余或虚假的替代方案,并验证每个子节点在其父节点之外是否仍有额外贡献。由此在SAE特征上诱导出的拓扑结构定义了一个可微的结构损失,用于引导SAE训练,而拓扑引导的细化过程则缓解了特征吸收问题,并利用学习结构所暴露出的持续重建差距来初始化新特征。随后,通过重新评估每个特征的完整父节点集合,从修订后的字典中再次诱导出整个图,从而闭合字典-图自一致性循环。实验表明,在受控玩具模型中实现了精确的混合拓扑恢复,在真实LLM激活上相比结构化和事后基线具有更高的关联可靠性和语义有效性,并且相比传统SAE特征展现出更强的特征级因果干预能力。因此,AG-SAE将恢复出的混合拓扑特征结构转化为无监督训练信号,改进了字典,实现了超越树拓扑限制的可靠特征组织,并展现出超越重建的更强因果控制能力。
英文摘要
Sparse autoencoders (SAEs) expose interpretable features in large language model activations, yet existing structured SAEs impose single-parent trees or forests, while post-hoc graphs permit multiple parents but neither guide feature learning nor ensure reliable relation recovery. We introduce the Adaptive Graph Sparse Autoencoder (AG-SAE), a structure-guided training paradigm that treats each feature's complete parent set as an atomic structural hypothesis and lets evidence select zero, one, or multiple parents. By competing complete parent sets against null, subset, and alternative explanations, AG-SAE identifies jointly necessary multi-parent relations while rejecting redundant or spurious alternatives and verifying that each child contributes beyond its parents. The induced topology over SAE features then defines a differentiable structural loss that guides SAE training, while topology-guided refinement mitigates feature absorption and uses persistent reconstruction gaps exposed by the learned structure to initialize new features. The entire graph is then induced again from the revised dictionary by reassessing every feature's complete parent set, closing the dictionary-graph self-consistency cycle. Experiments demonstrate exact mixed-topology recovery in a controlled toy model, greater relational reliability and semantic validity than structured and post-hoc baselines on real LLM activations, and stronger feature-level causal interventions than conventional SAE features. AG-SAE thereby turns recovered mixed-topology feature structure into an unsupervised training signal that improves the dictionary, enables reliable feature organization beyond the topological limitations of trees, and exhibits stronger causal control beyond reconstruction.
发表机构
- University of Arizona(亚利桑那大学)
- Iowa State University(爱荷华州立大学)
- Old Dominion University(欧道明大学)
机构由 AI 辅助整理,请以论文原文为准。