AI 中文总结
本研究系统探究模型编辑用于安全代码生成,提出SafeEdit方法缓解模型编辑的安全与功能权衡,其与CoSec结合可提升安全代码生成性能。
AI 中文摘要
大型语言模型(LLMs)被广泛用于代码生成,但它们可能重现从不安全训练模式中学习到的有漏洞实现。现有工作主要探索推理时加固,即在不修改目标模型的情况下减少不安全生成,但依赖辅助组件并增加运行时开销。我们首次对模型编辑作为安全代码生成的模型级加固机制展开系统研究。我们评估了3种最先进的编辑方法,涵盖不同LLM家族,并将它们与代表性推理时方法CoSec对比,重点关注安全性、鲁棒性、泛化性和功能正确性。模型编辑在已见漏洞类型上的安全增益大于CoSec,较原始模型提升安全比率15%-25%,且在提示扰动下增益保持稳定。然而,这些改进向未见漏洞的转移不可靠,还可能降低功能正确性。为缓解该权衡,我们提出SafeEdit,一种结合功能调优与编辑感知正则化的后编辑优化方法。在8个目标LLM上,SafeEdit在T=0.1/0.4/0.8时较UltraEdit提升Pass@1 11.73/13.70/15.50个百分点,同时基本保留安全性;与CoSec相比,实现7.54%-12.04%的相对安全比率增益。在CodeGuard+上的额外评估证实,SafeEdit可提升安全且正确的生成效果。SafeEdit与CoSec互补,二者结合可在保持强功能正确性的同时进一步提升安全性。总体而言,我们的结果为将模型编辑应用于安全代码生成提供了有实证支持的指导。
英文摘要
Large language models (LLMs) are widely used for code generation, yet they can reproduce vulnerable implementations learned from insecure training patterns. Prior work has mainly explored inference-time hardening, which reduces insecure generations without modifying the target model but relies on auxiliary components and adds runtime overhead. We conduct the first systematic study of model editing as a model-level hardening mechanism for secure code generation. We evaluate 3 state-of-the-art editing methods across diverse LLM families and compare them with CoSec, a representative inference-time approach, focusing on security, robustness, generalization, and functional correctness. Model editing yields larger security gains than CoSec on seen vulnerability types, improving security ratios by 15%-25% over vanilla models, with gains remaining stable under prompt perturbations. However, these improvements transfer unreliably to unseen vulnerabilities and can reduce functional correctness. To mitigate this trade-off, we propose SafeEdit, a post-edit refinement method combining functional tuning with edit-aware regularization. Across eight target LLMs, SafeEdit improves Pass@1 over UltraEdit by 11.73/13.70/15.50 percentage points at T=0.1/0.4/0.8 while largely preserving security. Compared with CoSec, it achieves relative security-ratio gains of 7.54%-12.04%. Additional evaluation on CodeGuard+ confirms improved joint secure-and-correct generation. SafeEdit and CoSec are also complementary, and their combination can further improve security while maintaining strong functional correctness. Overall, our results provide evidence-backed guidance for applying model editing to secure code generation.
CommentsISSTA 2026