发表机构
Jagiellonian University; IDEAS Research Institute(雅盖隆大学; IDEAS研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ClasSAE通过可微特征-类别亲和力矩阵,在训练中引导稀疏自编码器实现类别对齐与自动标注,无需事后探测,并在ImageNet上验证了其有效性与改进的类别分离。
AI 中文摘要
稀疏自编码器(SAEs)最初是一种无监督工具,用于将神经表示分解为稀疏、可解释的特征,并且越来越多地不仅用于被动分析,还用于主动干预,如遗忘、偏见缓解和概念编辑。这些编辑和引导方法的一个核心挑战是可靠地将特征与目标概念匹配;目前大多数方法通过在已训练、冻结的字典上计算事后分数来解决这一问题。我们则引入了ClasSAE,一种新颖的方法,它既自动为特征分配类别,又在训练过程中引导编码器朝向类别可分离的表示。具体来说,我们对一个可训练的特征-类别亲和力矩阵应用可微的top-$k$算子,并带有每特征预算,将每个样本选择的特征与它们被训练来表示的类别耦合起来。由于梯度通过激活特征的选择而非仅通过其幅度流动,编码器和亲和力矩阵共同适应,而不是在分离的阶段中拟合。结果是得到一个既类别可分离又类别标注的字典,无需事后探测。我们提出了在该框架内强制稀疏性的三种变体,它们在整体性能上相当,但略有不同的权衡。使用ImageNet上的CLIP ViT-L/14嵌入,我们表明学习到的亲和力矩阵与在保留数据上独立估计的事后特征-类别矩阵非常一致。该模型还支持仅从编码器和亲和力矩阵直接进行类别预测,无需单独拟合的分类器,并且其更类别对齐的编码器在目标探针扰动评估中产生了改进的分离。此https URL。
英文摘要
Sparse Autoencoders (SAEs) began as an unsupervised tool for decomposing neural representations into sparse, interpretable features, and are increasingly used not only for passive analysis but also for active interventions such as unlearning, bias mitigation, and concept editing. A central challenge for these editing and steering methods is reliably matching features to target concepts; most current approaches address this by computing post-hoc scores over an already-trained, frozen dictionary. We instead introduce ClasSAE, a novel method that both automatically assigns classes to features and guides the encoder toward class-separable representations during training. Specifically, we apply a differentiable top-$k$ operator to a trainable feature--class affinity matrix with per-feature budgets, coupling the features selected for each sample to the classes they are trained to represent. Because gradients flow through the selection of active features rather than only through their magnitudes, the encoder and the affinity matrix co-adapt rather than being fit in separate stages. The result is a dictionary that is both class-separable and class-annotated, with no need for post-hoc probing. We propose three variants for enforcing sparsity within this framework, which achieve comparable overall performance with slightly different trade-offs. Using CLIP ViT-L/14 embeddings on ImageNet, we show that the learned affinity matrix agrees closely with an independently estimated post-hoc feature-class matrix computed on held-out data. The model also supports direct class prediction from the encoder and affinity matrix alone, without a separately fitted classifier, and its more class-aligned encoder yields improved separation in Targeted Probe Perturbation evaluations. https://github.com/St0pien/ClasSAE.