发表机构
University of Wisconsin–Madison; Socket Inc.(威斯康星大学麦迪逊分校; Socket公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CASHEWS是一个JavaScript预处理器,通过去混淆、提取模块和计算反向切片来精简代码,提高LLM恶意包检测覆盖率并降低假阴性率,使大规模分析更实用。
AI 中文摘要
恶意npm包检测工具现在利用LLM对源代码的语义理解,在大规模范围内检测恶意意图。这一能力在识别近期供应链攻击(如Shai-Hulud)中涉及的包方面已被证明具有不可估量的价值。然而,威胁行为者利用LLM有限的上下文窗口,通过JavaScript技术(如产生高令牌密度的代码混淆以及将恶意代码与良性包捆绑)来逃避检测,导致检测器跳过大型文件或遗漏恶意行为。这为规避检测创造了攻击面。在本文中,我们提出了CASHEWS,一个JavaScript预处理器,通过重写源代码以移除与分析无关或可能误导模型的代码来减小文件大小。给定一个包源文件,CASHEWS通过迭代解码对其进行去混淆,提取捆绑模块和动态执行的代码,识别恶意汇聚点并计算到达这些汇聚点的反向切片,并缩写长字面量和标识符,为检测器生成紧凑表示。在512个大型包文件、两种扫描器类型和三个LLM上,CASHEWS将分析覆盖率从69.1%–85.7%提高到98.8%–100%,并将假阴性率降低最多18.6个百分点。CASHEWS的中位预处理时间为30秒,同时将净分析成本降低34.6%,使得注册表范围的基于LLM的分析更加实用。通过在分析前预处理源代码,CASHEWS使研究人员和行业从业者能够以与较弱模型相同或更低的分析成本,使用更强大的模型进行恶意包检测。
英文摘要
Malicious npm package detection tools now leverage LLMs' semantic understanding of source code to detect malicious intent at scale. This capability has proven invaluable in identifying packages involved in recent supply-chain attacks such as Shai-Hulud. However, threat actors exploit the limited context windows of LLMs through JavaScript techniques such as code obfuscation that yields high token density and bundling malicious code with benign packages, causing detectors to skip large files or miss malicious behavior. This creates an attack surface for evading detection. In this paper, we present CASHEWS, a JavaScript preprocessor that reduces file size by rewriting source code to remove code that is irrelevant to analysis or likely to mislead the model. Given a package source file, CASHEWS deobfuscates it through iterative decoding, extracts bundled modules and dynamically executed code, identifies malicious sinks and computes backward slices that reach them, and abbreviates long literals and identifiers to produce a compact representation for the detector. Across 512 large package files, two scanner types, and three LLMs, CASHEWS increases analysis coverage from 69.1--85.7% to 98.8--100% and reduces the false-negative rate by up to 18.6 percentage points. CASHEWS also has a median preprocessing time of 30 seconds while reducing net analysis cost by 34.6%, making registry-wide LLM-based analysis more practical. By preprocessing source code before analysis, CASHEWS enables researchers and industry practitioners to use more powerful models for malicious package detection at the same or lower analysis cost as less powerful models.