arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

设备端语言模型安全性有多脆弱?用于稀疏故障分析的安全关键参数定位

How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis

Muhammad Zeeshan Karamat, Christiana Chamon Garcia

arXiv 2610.09000首次发表:更新:

发表机构

Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究定位LLaMA-2-7B-Chat中安全关键参数,发现仅修改0.19%权重即可显著提升攻击成功率,为设备端模型提供针对性防护依据。

AI 中文摘要

随着小型语言模型(SLM)越来越多地部署在资源受限和设备端平台上,包括作为智能体系统的组成部分,本地存储模型参数的完整性成为一个重要的安全问题。我们研究了LLaMA-2-7B-Chat中安全敏感行为是否集中在稀疏的参数子集内,从而为针对性分析创造更小的故障面。我们研究了两种互补的定位方法:低秩安全关联子空间分析和参数级安全-效用重要性过滤。两种方法都揭示了网络中高度非均匀的安全敏感性,其中MLP的down_proj始终作为突出的安全敏感组件出现,而o_proj的贡献较小。通过参数级定位,仅修改down_proj中0.19%的模型权重即可实现53%的基础ASR和56%的GCG ASR,而tinyBenchmarks准确率保持在51.6%,相比之下未修改的基线为52.2%。这些结果促使在资源受限、设备端和智能体环境中部署的语言模型进行针对性故障分析和选择性完整性保护。

英文摘要

As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.

CommentsAccepted at NeurIPS 2026 Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑