AI 中文总结
研究基于大语言模型的编码工具生成的自动化脚本安全风险,通过跨模型收集代码并审查漏洞,发现风险与任务类别有关,各模型漏洞情况相似,组织应思考是否不经审查就部署此类自动化代码。
AI 中文摘要
基于大语言模型的编码工具使非专业用户能够生成常规自动化脚本,这些脚本可能未经严格安全审查就进入企业工作流程。本研究直接考察此风险。在三个自动化领域使用相同提示从ChatGPT、Microsoft Copilot和Google Gemini收集代码,让Claude Code进行标准化漏洞审查。用CVSS v3.1对每个识别出的漏洞评分并映射到OWASP Top 10:2021和MITRE ATT&CK框架。每个脚本都包含可利用漏洞。17个识别出的漏洞类别中,9个出现在所有三个模型的代码中,14个出现在至少两个模型中。跨平台加权CVSS分数差异小于10%。风险并非与特定模型相关,而是与任务类别有关。因此,组织不应问信任哪个工具,而应问是否应不经审查就部署大语言模型生成的自动化代码。
英文摘要
LLM-based coding tools enable non-expert users to generate routine automation scripts that may enter enterprise workflows without meaningful security review. This study examines that risk directly. Code was collected from ChatGPT, Microsoft Copilot, and Google Gemini using identical prompts across three automation domains. Claude Code performed a standardized vulnerability review. Each identified vulnerability was scored using CVSS v3.1 and mapped to the OWASP Top 10:2021 and the MITRE ATT&CK frameworks. Every script contained exploitable vulnerabilities. Nine of the 17 identified vulnerability classes appeared in code from all three models, while 14 of the 17 vulnerability classes appeared in at least two models. The weighted CVSS scores across platforms differed by less than 10%. The risk is not tied to any particular model but rather to the task category. Organizations should therefore ask not which tool to trust, but instead whether LLM-generated automation code should be deployed without review.
Comments9 pages, 9 figures, 2 tables