发表机构
Thapar Institute of Engineering and Technology; University of Waterloo(塔帕尔工程技术学院; 滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型集成到软件开发中自主生成安全认证代码的能力,通过双模式评估框架结合四种提示策略评估五个AI编码助手,发现当前助手不能默认生成安全应用,企业部署需转向持续、标准驱动的验证管道。
AI 中文摘要
大语言模型(LLMs)越来越多地集成到软件开发工作流程中,但其自主生成安全认证代码的能力仍不确定。本文通过结合静态代码分析和动态渗透测试的双模式评估框架,根据NIST SP 800-63B指南评估了五个著名AI编码助手生成的认证系统的安全架构。研究了四种提示策略(基本、安全、基于NIST和重新提示)下模型的行为。实证结果表明,从功能或一般安全提示生成的代码始终缺少关键保护。虽然提供明确的单次NIST上下文可显著提高合规性,但仍存在结构上的不足。相反,需要迭代重新提示,迫使模型进入上下文自我审核循环,以实现全面、深度防御的安全架构。最终证明当前AI编码助手不能默认生成安全应用,企业部署必须从单次提示工程过渡到持续的、标准驱动的验证管道。
英文摘要
Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their ability to autonomously generate secure authentication code remains uncertain. This paper evaluates the security architecture of authentication systems generated by five prominent AI coding assistants through a bi-modal assessment framework combining static code analysis and dynamic penetration testing, mapped to NIST SP 800-63B guidelines. The study examines model behavior across four prompting strategies Basic, Secure, NIST-Based, and Reprompting to reflect varying levels of developer guidance. Empirical results demonstrate that code generated from functional or generically secure prompts consistently omits critical protections, particularly concerning brute-force resistance, session management, and robust password handling. While providing explicit, single-shot NIST context significantly improves compliance, the findings reveal that this remains structurally inadequate. Instead, iterative Reprompting: forcing models into a contextual self-auditing loop is strictly required to achieve a comprehensive, defense-in-depth security architecture. Ultimately, this study proves that current AI coding assistants do not produce secure-by-default applications, dictating that enterprise deployments must transition from single-shot prompt engineering to continuous, standards-driven verification pipelines.