发表机构
ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM部署时安全检测依赖监督数据的问题,本文基于线性表示假设和局部稀疏性,提出局部掩码SAE异常检测框架,仅用1-2%神经元即可近最优检测。
AI 中文摘要
大语言模型(LLM)的部署时安全方法主要采用监督式,并假设可以访问不安全的训练数据。然而,新的攻击和危害类别经常出现,这类监督式训练的模型无法捕获。另一种方法是通过异常检测的视角来看待这个问题,即仅依赖对安全数据建模并标记分布外输入。但是,LLM激活位于高维空间中,这引发了关于异常检测在统计上是否可行的担忧。我们表明,在线性表示假设(LRH)下,确实可能存在希望。在LRH概念空间中,通常通过稀疏自编码器(SAE)恢复,邻近点共享一个小的共同活跃支持集。利用这一局部稀疏性洞察,我们提出了一个基于局部掩码SAE的异常检测框架,并提供了理论依据。我们在各种架构和数据集上验证了它,包括能力测试数据集和安全特定数据集。最后,当我们允许算法使用1%的分布外数据进行校准,局部稀疏方法实现了接近最优的性能,展示了它们仅使用1-2%的SAE神经元进行计算就能捕获有意义的安全信息的能力。
英文摘要
Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Using this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications. We validate it on various architectures and datasets, including both capability-testing datasets and safety-specific datasets. Finally, when we allow algorithms to use 1% out-of-distribution data for calibration, locally sparse methods achieve near-optimal performance, demonstrating their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.
CommentsPublished at NeurIPS2026