arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CS-Guard:面向代码生成安全的LLM护栏基准测试

CS-Guard: Benchmarking LLM Guardrails for Code Generation Security

Jinyang Li, Mingyu Guo, Hung X. Nguyen

arXiv 2609.09798首次发表:更新:

发表机构

Adelaide University(阿德莱德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CS-Guard是首个系统评估代码生成安全护栏的基准,涵盖文本到代码和代码到代码任务,发现现有护栏在越狱和虚构场景攻击下成功率极高,并发布模块化评估框架。

AI 中文摘要

大型语言模型(LLM)已被利用来生成恶意软件,但针对代码生成安全的护栏有效性仍不明确。我们提出了CS-Guard,这是首个系统评估代码生成安全护栏的基准。它涵盖1)文本到代码生成,包含1000个高质量恶意软件生成提示、7种越狱攻击,以及一种新颖的虚构场景攻击(FSA),该攻击将恶意意图嵌入合法的虚构软件开发场景中;以及2)代码到代码生成,包含331个代码提示,涵盖代码填充、代码补全和代码翻译。我们在七个LLM上实证评估了9种护栏。我们发现,当前护栏在应对恶意代码生成请求时表现不佳:对于文本到代码,许多护栏在越狱后的平均攻击成功率(ASR)达到约50%;对于代码到代码,基础LLM上的平均ASR接近100%,且在许多护栏中仍保持较高水平(14.4%至接近100%)。我们的FSA在许多护栏上也实现了接近100%的ASR,这对现实世界软件开发提出了重大可靠性担忧。为支持未来研究,CS-Guard采用模块化三层护栏分类法,允许开发者注册护栏进行评估。我们发布了该基准和数据,以便社区进一步评估。

英文摘要

Large language models (LLMs) have been ex- ploited to generate malware, but the effective- ness of guardrails for code generation secu- rity remains unclear. We introduce CS-Guard, the first benchmark to systematically evalu- ate guardrails for code generation security. It covers 1) text-to-code generation with 1000 high-quality malware-generation prompts, 7 jailbreak attacks, and a novel fictional scenario attack (FSA) that embeds malicious intent in a legitimate fictional software-development sce- nario; and 2) code-to-code generation with 331 code prompts spanning code infilling, code completion, and code translation. We empiri- cally evaluate 9 guardrails across seven LLMs. We find that current guardrails perform poorly against malicious code-generation re- quests: for text-to-code, the average attack success rate (ASR) after jailbreaks reaches about 50% for many guardrails; for code-to- code, average ASR approaches 100% on base LLMs and remains high across many guardrails (14.4% to nearly 100%). Our FSA also achieves ASR close to 100% across many guardrails, raising major reliability concerns for real-world software development. To sup- port future research, CS-Guard uses a modular three-layer guardrail taxonomy that lets devel- opers register guardrails for evaluation. We release the benchmark and data to enable fur- ther community evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑