arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30025cs.AI

面向安全且正确的代码生成的解释与引导

Interpreting and Steering for Safe and Correct Code Generation

  • George Mason University(乔治梅森大学)

机构由 AI 辅助整理,请以论文原文为准。

Hao Yan, Ziyu Yao

中文总结 AI 辅助

本研究通过CodeSec-Pairs数据集解释LLMs代码生成的安全机制,提出DuoSteer双引导方法,在5类漏洞实验中实现26.9%漏洞率降低,效果优于基线方法且可跨模型复现。

中文摘要 AI 辅助

大型语言模型(LLMs)经常生成包含漏洞的源代码,但很少有研究探讨区分安全代码与漏洞代码生成的内部机制。在本研究中,我们对LLMs进行了系统性的机制解释,旨在理解代码安全与漏洞的表征或驱动方式,以及如何将这些见解转化为可操作的引导策略,以鼓励生成更安全的代码。为此,我们引入了CodeSec-Pairs数据集,该数据集包含9342对Python安全代码与漏洞代码的对比样本,采样自Llama-3.1-8B-Instruct模型。利用该数据集,我们探索了定位与代码安全相关的层和注意力头的方法,并进一步实验了多种用于推理时减少漏洞的引导策略。我们提出了DuoSteer,一种双引导方法,同时对注意力头应用安全引导和代码正确性引导。在针对5种漏洞类型的实验中,DuoSteer实现了平均26.9%的漏洞率降低和7.5%的功能正确性提升,其表现不仅优于其他引导变体,也优于提示工程和监督微调基线方法。该优势在Qwen-2.5-Coder-7B-Instruct模型上也得到了复现,我们从该模型中额外采样了2500对对比样本用于验证。

英文摘要

Large language models (LLMs) frequently generate source code containing vulnerabilities, yet little work studies the internal mechanisms that distinguish safe from vulnerable generation in them. In this work, we systematically perform a mechanistic interpretation of LLMs, aiming at both understanding how code safety-vs-vulnerability is represented or driven by components in an LM and turning the insights into actionable steering strategies to encourage safer code generation. To this end, we introduce CodeSec-Pairs, a dataset of 9,342 Python safe-and-vulnerable contrastive code pairs, sampled from Llama-3.1-8B-Instruct. Utilizing the dataset, we explore approaches to localize layers and attention heads that relate to code safety, and further experiment with different steering strategies for inference-time vulnerability reduction. In particular, we propose DuoSteer, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads. In experiments over five vulnerability types, DuoSteer leads to an average of -26.9% vulnerability rate reduction and +7.5% functional correctness improvement, which outperforms not only other steering variants but also prompting and supervised fine-tuning baselines. The advantage also replicates on Qwen-2.5-Coder-7B-Instruct with another 2,500 contrastive pairs sampled from that model.

补充信息

↑