arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26222cs.LGcs.AIcs.CRcs.SE

NeuronFuzz:面向大语言模型安全评估的安全神经元引导模糊测试

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner, Sana Belguith, Lichao Wu

首次发表
浏览论文内容

中文总结 AI 辅助

NeuronFuzz是一种安全神经元引导的LLM安全评估白盒模糊测试框架,通过利用安全神经元激活值生成反馈,在21个模型上实现了高越狱发现率和零样本迁移能力,性能优于基线方法。

中文摘要 AI 辅助

安全评估对于判断对齐后的大语言模型(LLMs)是否仍能抵御越狱攻击至关重要。然而,现有的自动化测试方法大多依赖响应级反馈:每个候选提示通常需要生成目标模型的响应以评估其攻击有效性,该过程成本高昂,且更重要的是,对于强对齐模型仅能提供稀疏指导,其中大部分候选提示会以相同的失败结果被拒绝。本文提出NeuronFuzz,一种利用内部安全神经元作为大语言模型安全评估连续执行反馈的白盒模糊测试框架。SafetyOracle将安全神经元激活值转换为连续的安全告警分数,作为模糊测试的反馈,且可在预填充阶段获取,消除了模糊测试循环中的响应生成。为构建SafetyOracle,NeuronFuzz使用与模板无关的有害和良性输入以及感知稳定性的选择方法,识别出一组紧凑的安全神经元,其激活值可捕捉有害意图识别。此外,由于安全告警分数是可微的,NeuronFuzz利用其梯度识别安全敏感的模板位置,并使用掩码语言模型生成流畅、上下文兼容的变异,同时保留原始有害有效载荷并避免额外优化变量。我们在21个文本和多模态模型上评估NeuronFuzz,在5个白盒源模型上,它实现了76%-100%的越狱发现率,比基线方法高出多达48个百分点;其优化后的模板进一步零样本迁移到开放权重和6个专有目标模型,分别实现了69.6%/92.6%的平均攻击成功率(ASR)和前5集成攻击成功率(EASR),以及44.1%/60.0%的对应值。

英文摘要

Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to evaluate its attack effectiveness. This process is expensive and, more importantly, provides only sparse guidance on strongly aligned models, where most candidates are rejected with the same failure outcome. This paper presents NeuronFuzz, a white-box fuzzing framework that exploits internal safety neurons as continuous execution feedback for LLM safety evaluation. A SafetyOracle converts safety-neuron activations into a continuous safety alarm score that serves as feedback for fuzzing and can be obtained during prefill, eliminating response generation from the fuzzing loop. To construct the SafetyOracle, NeuronFuzz uses template-invariant harmful and benign inputs and stability-aware selection to identify a compact set of safety neurons whose activations capture harmful-intent recognition. Moreover, since the safety alarm score is differentiable, NeuronFuzz uses its gradients to identify safety-sensitive template positions and a masked language model to generate fluent, context-compatible mutations while preserving original harmful payload and avoiding additional optimization variables. We evaluate NeuronFuzz across 21 text and multimodal models. Across five white-box source models, it achieves a 76-100% jailbreak discovery rate, outperforming baselines by up to 48 percentage points. Its optimized templates further transfer zero-shot to open-weight and six proprietary target models, achieving average ASR and top-5 ensemble ASR (EASR) of 69.6%/92.6% and 44.1%/60.0%, respectively.

发表机构

  • University of Bristol(布里斯托尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑