arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过专家隔离与关闭实现大语言模型后门遏制

Backdoor Containment via Expert Quarantine and Shutdown in LLMs

Jianwei Li, Min-Seon Kim, Jung-Eun Kim

arXiv 2610.00663首次发表:更新:

发表机构

North Carolina State University(北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出“学习但引导”策略,通过隔离专家关闭(QES)在训练中引导后门进入可禁用组件,部署时零权重即可遏制,ASR降至0-10%。

AI 中文摘要

带有后门的大语言模型(LLMs)在良性输入上表现正常,但在隐藏触发器下会产生攻击者指定的输出。现有防御措施涵盖四个阶段——训练前、训练中、训练后和推理时——并共享两种基本策略之一:要么抑制后门学习(通过过滤有毒数据或在优化过程中中断其获取),要么先学习后净化(在完全后门模型形成后修复模型权重或门控输入)。我们提出第三种策略:学习但引导——允许训练期间形成后门,但将其路由到指定的、可隔离的组件中,该组件可在部署时禁用。为此,我们提出了隔离专家关闭(QES),一种在正则化引导的类MoE设置中构建的计算高效遏制策略。具体来说,给定一个有毒数据集,QES通过路由的专家特定LoRA分支和轻量级路由器增强基于Transformer的语言模型,并使用辅助路由目标将触发条件行为吸引到指定专家中,同时在其他地方保留良性能力。在部署时,缓解简化为单个常数时间操作:将隔离专家的路由权重置零,无需触发器筛选或进一步更新模型权重。实验上,我们的方法在两个任务、三种攻击和四个模型家族的大多数设置中将攻击成功率(ASR)从100%降至0-10%,同时下游效用通常得以保留或仅受到适度影响。这些结果确立了“学习但引导”作为生成式LLMs中后门遏制的一个先前未探索的领域。

英文摘要

Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert's routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate ASR from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish learn, but channel as a previously unexplored regime for backdoor containment in generative LLMs.

CommentsNeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑