arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

后门稀疏自编码器

Backdooring Sparse Autoencoders

Enrico Ahlers, Daniel Passon, Tobias Kiecker, Eik Reichmann, Lars Grunske

arXiv 2610.06049首次发表:更新:

发表机构

Humboldt-Universität zu Berlin(柏林洪堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示稀疏自编码器(SAE)可被植入后门,在不改动语言模型的情况下诱导特定行为,证明SAE是安全敏感组件,需防范供应链攻击。

AI 中文摘要

稀疏自编码器(SAEs)不仅越来越多地用于解释语言模型,还用于干预其内部表示。我们表明,这创造了一个供应链攻击面:一个被恶意修改的SAE在插入到原本未改变的语言模型的前向传播中时,可以诱导攻击者选择的行为。我们引入了一种仅解码器的SAE后门,该后门使底层LLM和SAE编码器保持冻结,将攻击限制在单个插入层的单个辅助组件上。以代码生成为案例研究,我们展示了在三个语言模型和广泛的插入层中,未经请求的代码插入的高比率,以及由提示线索触发的依赖触发器的行为。我们进一步使用HumanEval和选定的SAEBench指标评估修改后的SAE。虽然攻击有效性因模型和层而异,但强后门行为可以与几种常规SAE质量指标的相对较小变化共存。这些结果表明,SAE可以在不修改语言模型本身的情况下携带行为后门,因此应被视为安全敏感组件。

英文摘要

Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM and the SAE encoder frozen, restricting the attack to a single auxiliary component at a single insertion layer. Using code generation as a case study, we demonstrate high rates of unsolicited code insertion across three language models and a wide range of insertion layers, as well as trigger-dependent behavior conditioned on a prompt cue. We further evaluate the modified SAEs using HumanEval and selected SAEBench metrics. While attack effectiveness varies across models and layers, strong backdoor behavior can coexist with relatively small changes in several conventional SAE quality measures. These results establish that SAEs can carry behavioral backdoors without modifying the language model itself and should therefore be treated as security-sensitive components.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑