arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过变形测试和关联规则挖掘对大语言模型生成代码进行横切安全分析

Cross-Cutting Security Analysis of LLM-Generated Code via Metamorphic Testing and Association Rule Mining

Zedong Peng, Chenggang Wang, Shangyue Zhu

arXiv 2607.12089首次发表:更新:

AI 中文总结

研究针对LLM生成代码的安全漏洞,提出结合变形测试与关联规则挖掘的框架,定义九个MRs并应用于3700个代码片段,揭示共同违反结构与横切模式,还进行提示级风险分析,发现不安全代码生成具结构化和提示依赖性,推动相关验证与干预。

AI 中文摘要

大语言模型(LLMs)经常生成存在安全漏洞的代码,这些弱点很少被孤立:它们常常同时跨越多个关注领域,反映了软件安全的横切性质。我们提出一个框架,将面向安全的变形关系(MRs)与关联规则(AR)挖掘相结合,以检测LLM生成代码中的漏洞,揭示它们的共同违反结构,并将该结构追溯到提示级风险因素。我们定义了涵盖主要CWE类别的九个MRs,包括SQL注入、XSS、命令注入、路径遍历、硬编码凭证、弱加密和内存安全错误,并使用基于LLM的评判器将其应用于来自LLMSecEval基准的五个开放模型生成的3700个代码片段。结果表明,68.8%的片段违反了至少一个MR,硬编码凭证(79.1%)和命令注入(74.4%)是最普遍适用的失败类型。AR挖掘揭示了强大的横切共同违反模式,特别是XSS和弱加密共同违反以82.5%的置信度(提升=3.23)预测硬编码凭证,以及将认证、凭证处理和加密弱点以及输入处理和内存安全失败联系起来的紧密耦合集群。然后我们进行提示级风险分析,发现与数据库和认证相关的提示是广泛横切不安全的强预测指标,而65.5%的提示在所有五个模型中产生一致的违反结果。这些发现表明,不安全的代码生成不仅仅是独立缺陷的集合,而是一种结构化的、由提示条件决定的现象,促使进行集群感知验证和提示级干预,以实现更安全的LLM辅助编程。

英文摘要

Large language models (LLMs) frequently generate code with security vulnerabilities, yet these weaknesses are rarely isolated: they often span multiple concern areas simultaneously, reflecting the cross-cutting nature of security in software. We present a framework that combines security-oriented Metamorphic Relations (MRs) with Association Rule (AR) mining to detect vulnerabilities in LLM-generated code, uncover their co-violation structure, and trace that structure back to prompt-level risk factors. We define nine MRs covering major CWE categories, including SQL injection, XSS, command injection, path traversal, hard-coded credentials, weak cryptography, and memory-safety errors, and apply them using an LLM-based judge to 3,700 code snippets generated by five open models from the LLMSecEval benchmark. The results show that 68.8% of snippets violate at least one MR, with hard-coded credentials (79.1%) and command injection (74.4%) among the most prevalent applicable failures. AR mining reveals strong cross-cutting co-violation patterns, notably that XSS and weak cryptography co-violations predict hard-coded credentials with 82.5% confidence (lift = 3.23), along with tightly coupled clusters linking authentication, credential handling, and cryptographic weakness, as well as input-handling and memory-safety failures. We then perform prompt-level risk analysis and find that database- and authentication-related prompts are strong predictors of broad cross-cutting insecurity, while 65.5% of prompts yield consistent violation outcomes across all five models. These findings show that insecure code generation is not merely a collection of independent defects, but a structured and prompt-conditioned phenomenon, motivating cluster-aware verification and prompt-level intervention for safer LLM-assisted programming.

CommentsThis work has already been accepted by IEEE 27th International Conference on Information Reuse and Integration for Data Science

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑