arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

土拨比特翻转攻击:通过比特翻转在混合专家大语言模型中植入无限生成循环

Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips

Huakang Lin, Tiancheng Zheng, Mingxuan Sun, Tianhong Xu, Fan Zhang, Yunsi Fei, Ruyi Ding

arXiv 2608.25276首次发表:更新:

发表机构

Louisiana State University; Northeastern University; Zhejiang University; University of California, Los Angeles; Southeast University(路易斯安那州立大学; 东北大学; 浙江大学; 加利福尼亚大学洛杉矶分校; 东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出首个针对混合专家大语言模型的土拨比特翻转攻击,通过翻转路由层相关比特,可使平均输出膨胀率达5912%,揭示了MoE架构的鲁棒性漏洞。

AI 中文摘要

混合专家(MoE)架构通过路由机制选择性激活专家子网络,实现了大语言模型(LLM)的可扩展性与高效性。但这种自适应设计引入了新的攻击面:特定专家与某些 token(如序列结束符)存在不成比例的关联,攻击者可通过轻量扰动操纵模型行为。本研究提出**土拨比特翻转攻击(GBFA)**,这是首个针对MoE类LLM的基于比特翻转的“钱包可用性拒绝攻击”。通过识别并翻转与相关专家激活关联的路由层比特,研究人员证明GBFA可在对话、推理、智能体任务三种不同LLM模式下大幅延长解码 token 用量,同时在很大程度上保持语义保真度。在四种主流现实世界MoE类LLM上,手动停用平均少于**4个专家**即可使平均输出膨胀率达到$\boldsymbol{5912\%}$,多数测试样本达到最大 token 数。这些结果揭示了MoE架构对比特翻转的鲁棒性漏洞,凸显了GBFA作为针对LLM的可用性攻击的潜力。

英文摘要

Mixture-of-Experts (MoE) architectures enable scalable and efficient large language models (LLMs) by selectively activating expert sub-networks through a routing mechanism. However, this adaptive design introduces a new attack surface: specific experts become disproportionately correlated with certain tokens (e.g., end-of-sequence), allowing adversaries to manipulate model behavior via lightweight perturbations. In this work, we present \textbf{Groundhog Bit-Flip Attack (GBFA)}, the first bit-flip-based \textit{ Denial-of-Wallet availability attack} against MoE-based LLMs. By identifying and flipping routing-layer bits associated with related expert activations, we demonstrate that GBFA substantially extends the decoding token usage across three different LLM modes: conversational, reasoning, and agentic tasks, while largely preserving semantic fidelity. Across four main real-world MoE-based LLMs, manually deactivating on average fewer than \textbf{4 experts} drives average output inflation to $\mathbf{5912\%}$, with the majority of test samples reaching max tokens. These results reveal a robustness vulnerability of MoE architectures to bit flip, and highlight the potential of GBFA as an availability attack against LLMs.

Comments9 pages, 3 figures; Accepted at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑