arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30403cs.CR

为何LLM后门防御呈现碎片化?基于稀疏自编码器的特征级解释

Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders

Yizhe Zeng, Chenxu Niu, Wei Zhang, Hao Huang, Yunpeng Li, Dongxu Han, Dan Du, Cheng Hong, Hequn Xian, Yuling Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究利用稀疏自编码器开展LLM后门的特征级机制分析,揭示脏标签与干净标签后门的编码差异,提出推理时特征钳位方法,可降低攻击成功率同时保留良性任务性能,解释了现有防御的碎片化问题。

中文摘要 AI 辅助

后门攻击对大型语言模型(LLMs)构成严重威胁,但现有防御手段仍呈碎片化,无法为脏标签和干净标签攻击提供统一防御。为探究这种碎片化产生的原因,我们首次利用稀疏自编码器(SAEs)对LLM后门开展系统的特征级机制分析。通过对干净和中毒模型在干净与触发输入上进行2×2对比,我们将后门诱导的logit偏移追溯至高贡献SAE特征,并将其分为四类角色:交互特征、被抑制特征、混合特征和权重修改特征。该分类揭示了系统的编码差异:脏标签后门由孤立的交互特征主导,而干净标签后门则更多依赖混合特征与权重修改特征的异质组合。这些差异解释了为何现有防御在不同攻击范式下呈现碎片化。我们通过推理时特征钳位验证了这一假设,该方法在多数脏标签设置下将攻击成功率(ASR)降至至多10.8%,在多数干净标签设置下降至至多15.4%,同时保留了良性任务性能。这些结果表明,基于SAE的分析可解释防御碎片化问题,并为可解释的后门缓解提供指导。

英文摘要

Backdoor attacks pose a serious threat to large language models (LLMs), but existing defenses remain fragmented, failing to pro?vide unified defense against both dirty-label and clean-label attacks. To investigate why such fragmentation arises, we present the first systematic feature-level mechanistic analysis of LLM backdoors using sparse autoencoders (SAEs). Starting from a 2 x 2 comparison of clean and poisoned models on clean and triggered inputs, we trace backdoor-induced logit shifts to high-contributing SAE features and categorize them into four roles: interac?tion, suppressed, mixed, and weight-modified features. This taxonomy reveals system?atic encoding differences: dirty-label back?doors are dominated by isolated interaction features, whereas clean-label backdoors rely more on heterogeneous mixtures of mixed and weight-modified features. These differ?ences explain why existing defenses remain fragmented across attack paradigms. We val?idate this hypothesis through inference-time feature clamping, which reduces ASR to at most 10.8% in most dirty-label settings and at most 15.4% in the majority of clean-label settings, while preserving benign-task perfor?mance. These results show that SAE-based analysis can explain defense fragmentation and guide interpretable backdoor mitigation.

发表机构

  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
  • School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • Ant Group(蚂蚁集团)
  • College of Computer Science and Technology, Qingdao University(青岛大学计算机科学与技术学院)
  • Institute of Cryptography and Cyber Security (Whampoa)(cryptography and cyber security institute)

机构由 AI 辅助整理,请以论文原文为准。

↑