arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07746cs.CL

LLM 取证:后门藏在哪里?利用稀疏自编码器定位和控制触发机制

LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders

  • Inria Paris(法国国家信息与自动化研究所巴黎分部)

机构由 AI 辅助整理,请以论文原文为准。

Wissam Antoun, Francis Kulumba, Théo Lasnier, Benoît Sagot, Djamé Seddah

AI总结:

本研究利用稀疏自编码器在受控语言切换场景中定位大语言模型后门的触发机制,发现检测、传播和语言跟踪由不同特征分工,残差流特征可控制行为。

AI中文摘要:

尽管大语言模型中的后门问题日益受到关注,但其内部工作机制仍受到严密审查。基于触发器的后门在行为上易于定义,即一种罕见的输入使模型切换到预设的响应模式,但触发器与其响应之间的机制尚不明确。我们在一个受控的、无害的语言切换场景中研究这一机制,其中固定的触发器序列使 1B 和 8B 语言模型以法语或德语继续生成英文提示。为此,我们跨层和跨 Transformer 组件训练稀疏自编码器(SAEs),然后将触发提示与翻译和预训练对照组进行比较,以识别与触发器相关的特征方向。我们展示了 SAE 特征如何以接近完美的 F1 分数将触发提示与对照组区分开来,但检测触发器的特征并不一定控制行为。在干预测试中,注意力层和 MLP 特征通常在触发提示上可靠激活,使其成为良好的检测器,但消融它们很少抑制语言切换,激活它们也很少诱导语言切换。相比之下,残差流特征在消融时可以抑制触发生成,且一些选定的特征可以在没有触发器的情况下诱导目标语言的续写。总之,这些令牌触发机制分解为不同的 SAE 特征方向,分别具有用于触发器检测、残差流传播和后续语言跟踪的独立特征。这种角色级分解最有可能迁移到其他基于触发器的后门,即使负载、层或电路位置不同。

英文摘要:

Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny. Trigger-based backdoors are easy to define behaviorally, a rare input that makes the model switch to a chosen response pattern, but the mechanism between triggers and their responses is less clear. We study this mechanism in a controlled, harmless language-switching setting, where fixed trigger sequences make 1B and 8B language models continue English prompts in French or German. For this, we train sparse autoencoders (SAEs) across layers and transformer components, then compare triggered prompts with translation and pretraining controls to identify trigger-relevant feature directions. We show how SAE features separate triggered prompts from controls with near-perfect F1, but features that detect the trigger do not necessarily control the behavior. In intervention tests, attention and MLP features often fire reliably on triggered prompts, making them good detectors, but ablating them rarely suppresses the language switch and activating them rarely induces it. In contrast, residual-stream features can suppress triggered generation when ablated, and some selected features can induce target-language continuations without the trigger. In short, these token-trigger mechanisms decompose into distinct SAE feature directions, with separate features for trigger detection, residual-stream propagation, and later language tracking. This role-level decomposition is the part most likely to transfer to other trigger-based backdoors, even when the payload, layers, or circuit locations differ.

补充信息

↑