arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JavaScript和TypeScript中易受攻击与已修复源代码的函数级数据集

A Function-level Dataset of Vulnerable and Fixed Source Code in JavaScript and TypeScript

Tamás Viszkok, Péter Hegedűs

arXiv 2609.38012首次发表:更新:

发表机构

University of Szeged(塞格德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对JavaScript和TypeScript的漏洞检测训练数据不足问题,构建了函数级数据集JsVul,通过语言特定流水线、多阶段去重和启发式标注确保质量,支持稳健模型训练。

AI 中文摘要

JavaScript和TypeScript在现代Web开发中被广泛使用,这使得它们的安全性至关重要;然而,自动化漏洞检测常常受到高质量训练数据可用性的限制。在此,我们提出了JsVul,一个从七个主要来源整理的数据集。与可能保留噪声(如压缩代码和外观性编辑)的通用多语言数据集不同,JsVul采用了一种特定于语言的流水线。我们收集了安全修复前后版本的文件,并通过过滤无关工件和应用自动语法规范化,隔离了与安全相关的更改。我们通过多阶段去重和基于启发式的标注确保了数据的完整性。JsVul以按时间排序的JSONL格式提供,支持JavaScript和TypeScript生态系统中稳健的模型训练,并展示了在构建漏洞数据集时语言感知预处理的重要性。

英文摘要

JavaScript and TypeScript are widely used in modern web development, making their security critical; however, automated vulnerability detection is often constrained by the availability of high-quality training data. Here we present JsVul, a dataset curated from seven major sources. Unlike generic multi-language datasets that may retain noise -- such as minified code and cosmetic edits -- JsVul utilizes a language-specific pipeline. We collected pre-fix and post-fix versions of files around security fixes and, by filtering irrelevant artifacts and applying automated syntax normalization, isolated security-related changes. We ensured data integrity through multi-stage deduplication and heuristic-based labeling. Provided in a time-ordered JSONL format, JsVul supports robust model training in the JavaScript and TypeScript ecosystem and demonstrates the importance of language-aware preprocessing in building vulnerability datasets.

CommentsThis manuscript is currently under review at Scientific Data

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑