arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MOMAT:多图谱混合用于低功耗量化大语言模型越狱防御

MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs

Boyang Li, Bingyu Shen, Weihao Hong, Zhiyuan Jiang, Xinlei Guan, Yan Ma, Miles Q. Li, Yi Sheng, Ruiyang Qin

arXiv 2610.01058首次发表:更新:

发表机构

Department of Computer Science and Engineering, University of Notre Dame(圣母大学计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MOMAT提出多图谱混合与存内计算加速的硬件增强框架,实现低功耗量化大语言模型越狱防御,兼顾安全性与能效。

AI 中文摘要

量化大语言模型因其低延迟和高能效而越来越多地部署在边缘设备上。然而,模型量化削弱了对齐保障,使量化大语言模型(qLLMs)极易受到越狱攻击。为解决这一挑战,我们提出了MOMAT(多图谱混合),一种硬件增强的安全框架,结合了结构化知识检索与低功耗防御加速。每个图谱代表有害或良性样本集和策略模板的语义簇,实现领域局部化的检索增强生成防护,缓解大型异构安全数据库中的维度灾难及由此产生的语义稀疏问题。MOMAT为每个提示从所有图谱中检索top-$k$相似特征,并使用轻量级MoE(专家混合)检测器进行评估,同时CiM(存内计算)加速的相似性引擎执行快速、低功耗的图谱局部检索。MOMAT基于CiM的检索将100个查询的批次从15,052.44毫秒加速至3,207.21纳秒($4.69 \times 10^6\times$加速),并将能量从$8.1 \times 10^7$ $\mu$J降至3.32 $\mu$J,相较于基于DRAM(树莓派)的基线实现了约$2.5 \times 10^5\times$的能耗降低。跨标准基准的红队评估表明,MOMAT在匹配最先进方法防御性能的同时,避免了良性过度杀伤并提供了显著的效率提升,证明基于CiM的模块化防御能使边缘部署的qLLMs更安全且更节能。我们将发布完整的223.2k样本数据集以促进未来研究。

英文摘要

Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address this challenge, we present MOMAT (Mixture of Multiple Atlases), a hardware-enhanced safety framework that combines structured knowledge retrieval with low-power defense acceleration. Each atlas represents a semantic cluster of harmful or benign sample sets and policy templates, enabling domain-localized Retrieval-Augmented Generation guarding that mitigates the curse of dimensionality and the resulting semantic sparsity problem in large, heterogeneous safety databases. MOMAT retrieves top-$k$ similarity features from all atlases for each prompt and evaluates them using a lightweight MoE (Mixture of Experts) detector, while a CiM (Compute-in-Memory)-accelerated similarity engine performs fast, low-power atlas-local retrieval. MOMAT's CiM-based retrieval accelerates a 100-query batch from 15,052.44 ms to 3,207.21 ns (a $4.69 \times 10^6\times$ speedup) and reduces energy from $8.1 \times 10^7$ $μ$J to 3.32 $μ$J, yielding an approximately $2.5 \times 10^5\times$ energy reduction over DRAM-based (Raspberry Pi) baselines. Red-team evaluations across standard benchmarks show that MOMAT matches the defense performance of state-of-the-art methods while avoiding benign overkill and providing substantial efficiency gains, demonstrating that CiM-based modular defenses can make edge-deployed qLLMs both safer and more energy-efficient. We will release the full 223.2k-sample dataset to foster future research.

Comments16 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑