arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpIn-ViT:设计一种具有机械可解释性的稀疏诱导视觉Transformer

SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable

Philip H. Lee, Parth Padalkar

arXiv 2608.14922首次发表:更新:

AI 中文总结

本研究提出端到端联合训练预训练ViT与改进SAE的SpIn-ViT框架,在9个图像分类基准上,其分类准确率、可解释性均优于现有事后SAE方法,还能生成更优的神经符号模型。

AI 中文摘要

机械可解释性最近已扩展到视觉Transformer(ViTs),稀疏自编码器(SAEs)越来越多地被用作事后工具,将内部表示分解为更稀疏、更可解释的特征。然而,由于事后SAEs是在ViT优化完成后对冻结的表示进行训练的,其潜在特征与下游分类目标并不直接对齐。我们引入SpIn-ViT,这是一个将预训练ViT与改进后的SAE进行端到端联合训练的框架,直接将稀疏的图像块级表示与图像分类对齐。SpIn-ViT学习语义一致的神经元激活,这些激活能定位有意义的图像区域,同时保持有竞争力的预测性能。我们在9个图像分类基准上,通过分类准确率、定量可解释性指标、基于AI的评估和人类评估对SpIn-ViT进行评估。与之前最先进的事后SAE方法相比,SpIn-ViT的平均分类准确率提高了8.84%,基于AI的可解释性评分几乎是其4倍,人类评估评分是其2倍以上。我们进一步使用SAE神经元提取可解释的规则集,创建神经符号模型,这些模型的平均分类准确率提高了5.97%,同时规则集规模比最先进的事后SAE方法创建的神经符号模型小58.8%。

英文摘要

Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post-hoc tools to decompose internal representations into sparse and more interpretable features. However, because post-hoc SAEs are trained on frozen representations after the ViT has already been optimized, their latent features are not directly aligned with the downstream classification objective. We introduce SpIn-ViT, a framework that jointly trains a pretrained ViT and a modified SAE end-to-end, directly aligning sparse patch-level representations with image classification. SpIn-ViT learns semantically coherent neuron activations that localize meaningful image regions while maintaining competitive predictive performance. We evaluate SpIn-ViT across nine image-classification benchmarks using classification accuracy, quantitative interpretability metrics, AI-based and Human evaluations. Compared with the previous state-of-the-art post-hoc SAE method, SpIn-ViT achieves 8.84% higher average classification accuracy, an AI-based interpretability score nearly four times as high, and a human-evaluation score more than twice as high. We further extract interpretable rule-sets using the SAE neurons to create neurosymbolic models which achieve 5.97% higher average classification accuracy while requiring a 58.8\% smaller rule-set than the neurosymbolic models created from the SOTA post-hoc SAE method.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑