为何LLM后门防御呈现碎片化?基于稀疏自编码器的特征级解释
Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders
浏览论文内容
中文总结 AI 辅助
该研究利用稀疏自编码器开展LLM后门的特征级机制分析,揭示脏标签与干净标签后门的编码差异,提出推理时特征钳位方法,可降低攻击成功率同时保留良性任务性能,解释了现有防御的碎片化问题。
中文摘要 AI 辅助
后门攻击对大型语言模型(LLMs)构成严重威胁,但现有防御手段仍呈碎片化,无法为脏标签和干净标签攻击提供统一防御。为探究这种碎片化产生的原因,我们首次利用稀疏自编码器(SAEs)对LLM后门开展系统的特征级机制分析。通过对干净和中毒模型在干净与触发输入上进行2×2对比,我们将后门诱导的logit偏移追溯至高贡献SAE特征,并将其分为四类角色:交互特征、被抑制特征、混合特征和权重修改特征。该分类揭示了系统的编码差异:脏标签后门由孤立的交互特征主导,而干净标签后门则更多依赖混合特征与权重修改特征的异质组合。这些差异解释了为何现有防御在不同攻击范式下呈现碎片化。我们通过推理时特征钳位验证了这一假设,该方法在多数脏标签设置下将攻击成功率(ASR)降至至多10.8%,在多数干净标签设置下降至至多15.4%,同时保留了良性任务性能。这些结果表明,基于SAE的分析可解释防御碎片化问题,并为可解释的后门缓解提供指导。
英文摘要
Backdoor attacks pose a serious threat to large language models (LLMs), but existing defenses remain fragmented, failing to pro?vide unified defense against both dirty-label and clean-label attacks. To investigate why such fragmentation arises, we present the first systematic feature-level mechanistic analysis of LLM backdoors using sparse autoencoders (SAEs). Starting from a 2 x 2 comparison of clean and poisoned models on clean and triggered inputs, we trace backdoor-induced logit shifts to high-contributing SAE features and categorize them into four roles: interac?tion, suppressed, mixed, and weight-modified features. This taxonomy reveals system?atic encoding differences: dirty-label back?doors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modified features. These differ?ences explain why existing defenses remain fragmented across attack paradigms. We val?idate this hypothesis through inference-time feature clamping, which reduces ASR to at most 10.8% in most dirty-label settings and at most 15.4% in the majority of clean-label settings, while preserving benign-task perfor?mance. These results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation.
发表机构
- Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
- School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Ant Group(蚂蚁集团)
- College of Computer Science and Technology, Qingdao University(青岛大学计算机科学与技术学院)
- Institute of Cryptography and Cyber Security (Whampoa)(cryptography and cyber security institute)
机构由 AI 辅助整理,请以论文原文为准。