arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12094cs.LGcs.AI

用于可解释的分布外检测的稀疏自编码器

Sparse Autoencoders for Interpretable Out-of-Distribution Detection

Ayush Karmacharya, Luke Luschwitz, Lucia Romero, Yanan Niu, Joseph Campbell

首次发表
浏览论文内容

中文总结 AI 辅助

研究如何可靠检测分布外样本,利用稀疏自编码器从神经网络中间激活中学习可解释特征,通过余弦相似度得出OOD分数,该方法在检测基准上性能领先且能提供可解释见解。

中文摘要 AI 辅助

可靠检测分布外(OOD)样本对机器学习模型的安全部署至关重要。神经网络对偏离训练数据的输入往往产生过度自信的预测,导致性能显著下降。许多OOD检测方法关注最终输出层,却忽略了中间网络层的丰富层次信息。本文引入一种利用稀疏自编码器(SAEs)从这些中间激活中学习可解释特征的新方法。发现分布内(ID)和OOD数据激活不同的稀疏特征集。提出一种基于测试样本稀疏特征激活与ID类平均激活之间余弦相似度的新OOD分数。我们的事后检测方法不仅在标准OOD检测基准上取得了领先性能,还对分布变化如何影响学习表示产生了可解释的见解。

英文摘要

Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models. Neural networks often produce overconfident predictions for inputs that deviate from their training data, leading to significant degradation in performance. While many OOD detection methods focus on the final output layer, they neglect the rich hierarchical information present in intermediate network layers. This paper introduces a novel approach that leverages sparse autoencoders (SAEs) to learn interpretable features from these intermediate activations. We find that in-distribution (ID) and OOD data activate distinct sets of these sparse features. We propose a new OOD score derived from the cosine similarity between the sparse feature activations of a test sample and the mean activations of ID classes. Our post-hoc detection method not only achieves state-of-the-art performance on standard OOD detection benchmarks, but yields interpretable insights into how distribution shift affects learned representations.

发表机构

  • Purdue University(普渡大学)
  • EPFL(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑