arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更新更大,但更安全?LLM生成代码功能-安全差距的纵向研究

Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code

Thiago Santos de Moura, Fynn Matuschek, Flavio Toffalini, Yannic Noller

arXiv 2610.08240首次发表:更新:

发表机构

Ruhr-Universität Bochum(波鸿鲁尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对32个LLM进行纵向分析,发现新模型绝对安全性提升但功能-安全差距未消除,开源家族显著缩小差距,紧凑模型通常更不安全,建议开发者勿假设升级或专有模型更安全。

AI 中文摘要

大型语言模型(LLM)被广泛用于生成代码。尽管其功能合理性持续提升,但生成的代码往往包含安全漏洞。功能-安全差距指的是通过功能测试但未通过安全测试的代码。近期一项针对三个模型家族的纵向研究得出结论:LLM变得更聪明但并未更安全,且唯一被考虑的开源权重家族停滞不前。这一结论是否适用于其他(开源权重)家族,尤其是紧凑型模型,仍属未知。我们针对来自七个模型家族(五个开源权重)的32个LLM进行了差距的纵向研究,覆盖每个家族连续三代旗舰版和紧凑版发布。使用CWEval中的119个任务(涵盖五种编程语言和31个CWE),我们比较了不同家族、模型规模和语言之间的轨迹。新模型在绝对意义上确实变得更安全,尽管没有家族完全消除差距。与先前工作不同,我们发现开源与否并不区分家族:每个被考虑的开源权重家族都显著缩小了差距,而Gemini 3.1 Pro保持的差距与Llama报告的差距一样大。紧凑模型通常产生比旗舰对应模型更不安全的代码,但有显著例外(如Gemini 3.7 Flash)。在CWE层面,我们确认了诸如日志注入(CWE-117)和HTTP响应拆分(CWE-113)等持续存在的弱点,以及最新专有模型在内存和整数弱点上的退化,并表明同一CWE在不同语言中承载的风险差异很大。从这些结果中,我们为LLM供应商、研究人员和开发者得出了启示。特别是,开发者不应假设升级会改善安全性,也不应假设专有模型更安全;他们应在每次模型变更后重新运行安全检查,并在模型上下文中提供安全的API。

英文摘要

Large Language Models (LLMs) are widely used to generate code. Although their functional plausibility keeps improving, the generated code often contains security vulnerabilities. The functionality-security gap captures code that passes functional tests but fails security tests. A recent longitudinal study of three model families concluded that LLMs become smarter but not safer, with the only considered open-weight family stagnating. Whether this holds for other (open-weight) families and particularly for compact models remains open. We present a longitudinal study of the gap across 32 LLMs from seven model families (five open-weight), covering three successive releases per family in flagship and compact variants. Using CWEval with 119 tasks in five programming languages and 31 CWEs, we compare trajectories across families, model sizes, and languages. Newer models do become safer in absolute terms, although no family closes the gap. Unlike prior work, we find that openness does not separate the families: every considered open-weight family narrows the gap significantly, while Gemini 3.1 Pro keeps a gap as wide as the one reported for Llama. Compact models usually produce less secure code than their flagship counterparts, with notable exceptions (e.g., Gemini 3.7 Flash). At the CWE level, we confirm persistent weaknesses such as log injection (CWE-117) and HTTP response splitting (CWE-113) and regressions in the newest proprietary models on memory and integer weaknesses, and show that the same CWE carries very different risk across languages. From these results, we derive implications for LLM vendors, researchers, and developers. In particular, developers should assume neither that upgrades improve security nor that proprietary models are more secure; they should rerun security checks after each model change and provide secure APIs in the model's context.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑