AI 中文总结
研究探索大语言模型中检测恶意代码的神经元,应用机械可解释性方法定位相关神经元,通过对恶意和良性PyPI包实验发现放大促进神经元、抑制抑制神经元可提升准确率,有助于了解模型编码恶意概念方式,为可靠防御机制提供思路。
AI 中文摘要
背景。大语言模型(LLMs)在理解和生成源代码方面能力日益增强,在软件工程任务中广泛应用。但其识别恶意或易受攻击代码模式的内部机制尚不清楚。目的。研究恶意软件检测行为在LLMs前馈网络(FFN)神经元中的编码位置,并通过因果干预验证归因,以识别检测恶意代码的最重要神经元。方法。应用机械可解释性方法在三个指令调整的LLMs中定位负责恶意软件检测行为的神经元。使用来自PyPI Malregistry的1500个恶意和1500个良性PyPI包,将行为归因于一组神经元。结果。实验结果表明,放大促进恶意软件检测的神经元同时抑制抑制性神经元可提高准确率,反之则使预测崩溃为单一类别,且程度和一致性高度依赖模型。结论。探索与安全相关知识的神经元有助于了解LLMs如何编码恶意编程概念,识别潜在有害的记忆行为,为更可靠的防御机制铺平道路。
英文摘要
Background. Large language models (LLMs) have become increasingly capable of understanding and generating source code, leading to their widespread adoption in software engineering tasks such as code completion, repair, and vulnerability detection. However, despite their strong empirical performance, the internal mechanisms through which LLMs recognize malicious or vulnerable code patterns remain poorly understood. Aim. We investigated where the malware detection behavior is encoded inside LLMs Feed Forward Network (FFN) neurons and verified the attribution with causal interventions on the neurons identified. This aims to identify the most important neurons in detecting malicious code. Methods. We applied mechanistic interpretability methods to locate the neurons being responsible for malware-detection behavior in three instruction-tuned LLMs: Llama3.1-8B-Instruct, Mistral-v0.3-7B-Instruct, and Qwen2.5-7B-Instruct. Using 1,500 malicious and 1,500 benign PyPI packages from the PyPI Malregistry, we attribute the behavior to a set of neurons. Results. The experimental results reveal that amplifying facilitating neurons for malware detection while suppressing inhibiting ones can boost accuracy, while the reverse collapses predictions toward a single class, although the magnitude and consistency is heavily model-dependent. We demonstrated that the guardrail detection mechanism varies across models, each represents its malware detection behavior differently within its FFN layers. Conclusions. Probing the neurons associated with security-relevant knowledge helps us gain insights into how LLMs encode malicious programming concepts, identify potentially harmful memorized behaviors, paving the way toward more reliable defense mechanisms, such as neuron-level editing, selective unlearning, and security-aware alignment for code-focused LLMs.
CommentsThe paper has been peer reviewed and accepted for publication in the 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)