AI 中文总结
研究大语言模型安全分类,利用稀疏自动编码器特征提取解决安全多面体中约束数量调整问题,在Qwen3.5 - 9B上对部分类别K = 2最优,准确率达96 - 99%,基于几何观点引入锥约束并经三阶段训练稳定。
AI 中文摘要
安全多面体(SaP)在大语言模型隐藏空间中学习线性半空间约束,但需要对约束数量K进行逐类别调整。我们表明稀疏自动编码器(SAE)特征提取解决了这一问题:在Qwen3.5 - 9B上,对于12/14个类别,K = 2成为最优值,在我们的BeaverTails分类基准上每个类别实现了96 - 99%的准确率,基本消除了详尽搜索的需要。这种向两个平面的收敛与线性表示假设一致。基于此几何观点,我们引入了一个锥约束,其可学习孔径适应每个类别的聚类集中度,并通过三阶段训练实现稳定。
英文摘要
Safety as Polytope (SaP) learns linear half-space constraints in LLM hidden space but requires per-category tuning of the constraint count K. We show that sparse autoencoder (SAE) feature extraction resolves this: K=2 becomes optimal for 12/14 categories on Qwen3.5-9B, achieving 96-99% accuracy per category on our BeaverTails classification benchmark, largely eliminating the need for exhaustive sweeps (K=4-25 with random initialization). This convergence to two planes is consistent with the Linear Representation Hypothesis, providing suggestive evidence that safety boundaries in this setting admit a low-dimensional linear description in the SAE feature space. Building on this geometric perspective, we introduce a cone constraint whose learnable aperture adapts to each category's cluster concentration, stabilized by a three-phase training