arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

海报:通过现实世界风险场景重新思考大语言模型代码生成中的安全性

Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios

Lixun Ma, Ruolong Ma, Bei Wang, Feng Wei, Zhenguang Liu, Lorenzo Cavallaro, Wentao Chen

arXiv 2607.23088首次发表:更新:

发表机构

Zhejiang University; University College London; MetaX Integrated Circuit Co., Ltd.; Artificial Intelligence Institute, CAICT(浙江大学; 伦敦大学学院; 元芯集成电路有限公司; 中国信息通信研究院人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究LLM代码生成的安全行为,识别出三种风险场景,构建含2700个测试用例的基准评估安全性,发现先进LLMs平均漏洞率超56%,证明安全感知提示可大幅降低风险。

AI 中文摘要

大语言模型(LLMs)广泛用于代码生成,但其在实际开发工作流程中的安全行为仍未得到充分探索。现有基准通常依赖明确指定的安全要求,无法捕捉提示往往模糊或不完整的现实世界场景。本文采用以开发者为中心的视角,识别出导致LLM生成代码中出现安全漏洞的三种代表性风险场景:模糊需求、操作上下文指定不足和安全与功能冲突。基于这些场景,构建了包含2700个测试用例的大规模基准,能在现实条件下对LLM安全性进行细粒度评估。对八个先进LLMs的广泛评估表明,所有模型在风险场景中的平均漏洞率超过56%。进一步证明安全感知提示可大幅降低这些风险,最多可提高45%。

英文摘要

Large Language Models (LLMs) are widely used for code generation, yet their security behavior in realistic development workflows remains underexplored. Existing benchmarks often rely on explicitly specified security requirements, failing to capture real-world scenarios where prompts are frequently ambiguous or incomplete. In this paper, we adopt a developer-centric perspective and identify three representative risk scenarios that commonly lead to security vulnerabilities in LLM-generated code: Ambiguous Requirements, Under-Specified Operational Context, and Security--Functionality Conflict. Based on these scenarios, we construct a large-scale benchmark comprising 2,700 test cases, enabling fine-grained evaluation of LLM security under realistic conditions. Extensive evaluation of eight state-of-the-art LLMs reveals that all models exhibit average vulnerability rates exceeding 56\% across risk scenarios. We further demonstrate that security-aware prompting can substantially mitigate these risks, achieving up to 45\% improvement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑