发表机构
Indraprastha Institute of Information Technology Delhi(德里英德拉普拉斯塔信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究实证评估大型语言模型在JavaScript代码漏洞识别中的效果,发现微调后准确率从29%提升至60%,优于传统SAST工具,可作补充手段。
AI 中文摘要
JavaScript 驱动着约 98.8% 的网站,因此其代码中的漏洞构成重大安全风险,然而现有的检测方法(如静态应用安全测试(SAST)工具)在应用于隔离代码片段时,往往无法识别许多真实世界的漏洞。本文对基于大型语言模型(LLM)的 JavaScript 程序漏洞识别进行了实证研究,评估了三个 LLM 系列(Gemini 1.5 Flash、GPT-4o Mini、DeepSeek-R1-Distill-Llama-8B),在包含 1,125 个 JavaScript 代码片段的数据集上,跨越五种常见弱点枚举(CWE)类别:注入(CWE-74)、操作系统命令注入(CWE-78)、跨站脚本(CWE-79)、SQL 注入(CWE-89)和不受控制的资源消耗(CWE-400),采用了多种提示策略(零样本、思维链、少样本)和微调方法。我们的实验表明,在片段级漏洞识别方面,LLM 显著优于传统 SAST 工具,微调的 Gemini 1.5 Flash 模型实现了 60% 的检测准确率,而基于规则的静态分析器性能接近零。我们发现,微调将准确率从 29% 提升至 60%,思维链提示有利于具备推理能力的模型(如 GPT-4o Mini),少样本提示对多态性漏洞(如跨站脚本)有效,且性能在不同漏洞类别间存在差异,对于结构化漏洞(如 SQL 注入)准确率可达 84%。这些结果表明,LLM 为 JavaScript 代码中的自动漏洞识别提供了一种实用方法,尤其是与任务对齐的监督相结合时,但由于召回率有限且不同漏洞类型间性能不均衡,它们应补充而非取代现有的安全分析工作流程。
英文摘要
JavaScript powers approximately 98.8% of all websites, making vulnerabilities in its code a significant security risk, yet existing detection approaches such as Static Application Security Testing (SAST) tools often fail to identify many real-world vulnerabilities when applied to isolated code snippets. This paper presents an empirical study of Large Language Model (LLM)-based vulnerability identification for JavaScript programs, evaluating three LLM families (Gemini 1.5 Flash, GPT-4o Mini, DeepSeek-R1-Distill-Llama-8B) across multiple prompting strategies (zero-shot, chain-of-thought, few-shot) and fine-tuning approaches on a dataset of 1,125 JavaScript code snippets spanning five Common Weakness Enumeration (CWE) categories: Injection (CWE-74), OS Command Injection (CWE-78), Cross-Site Scripting (CWE-79), SQL Injection (CWE-89), and Uncontrolled Resource Consumption (CWE-400). Our experiments show that LLMs substantially outperform traditional SAST tools on snippet-level vulnerability identification, with a fine-tuned Gemini 1.5 Flash model achieving 60% detection accuracy compared to near-zero performance from rule-based analyzers. We find that fine-tuning improves accuracy from 29% to 60%, Chain-of-Thought prompting benefits reasoning-capable models such as GPT-4o Mini, few-shot prompting is effective for polymorphic vulnerabilities such as Cross-Site Scripting, and performance varies across vulnerability categories, reaching up to 84% accuracy for structured vulnerabilities such as SQL Injection. These results indicate that LLMs provide a practical approach for automated vulnerability identification in JavaScript code, particularly when combined with task-aligned supervision, though they should complement rather than replace existing security analysis workflows due to limited recall and uneven performance across vulnerability types.
Comments8 pages, 1 figure, 4 tables