AI 中文总结
本文提出MalTotal框架,可跨5种主流语言检测恶意代码,F1值达93.1%,成本大幅降低,在12万个GitHub仓库中发现564个未知恶意仓库,适用于大规模多语言场景。
AI 中文摘要
开源软件(OSS)的广泛应用带来了重大安全风险,恶意代码投毒攻击日益成为公共包仓库和开源平台的攻击目标。现有检测方法包括基于启发式、学习及大语言模型(LLM)的方法,存在语言特定设计、泛化能力有限及分析成本高的问题,无法适配大规模多语言分析。为解决这些挑战,本文提出MalTotal,一种可扩展、高性价比的语言无关型恶意代码检测框架。MalTotal利用LLM辅助的语义推理识别敏感API,执行混合语义切片并重构恶意行为上下文,同时降低分析开销。评估结果显示,MalTotal在5种主流语言上的平均F1值达93.1%,优于8种最先进的基线方法;其混合切片将LLM令牌消耗降低94.0%,在2168个代码仓库上将分析成本从86.25美元降至5.19美元;在包含超730万个文件的12万个GitHub代码仓库大规模研究中,MalTotal以总计338美元的成本发现了564个此前未知的多语言恶意代码仓库。这些结果证明了MalTotal在缓解大规模代码投毒攻击方面的有效性、可扩展性及成本效益。
英文摘要
The widespread adoption of open source software (OSS) has introduced significant security risks, with malicious code poisoning attacks increasingly targeting public package registries and open-source platforms. Existing detection approaches, including heuristic-, learning-, and LLM-based methods, suffer from language-specific designs, limited generalization, and high analysis costs, making them unsuitable for large-scale multi-language analysis. To address these challenges, we propose MalTotal, a scalable and cost-effective framework for language-agnostic malicious code detection. MalTotal leverages LLM-assisted semantic reasoning to identify sensitive APIs, perform hybrid semantic slicing, and reconstruct malicious behavior contexts while reducing analysis overhead. Our evaluations show that MalTotal outperforms 8 state-of-the-art baselines, achieving an average F1-score of 93.1% across 5 mainstream languages. Its hybrid slicing reduces LLM token consumption by 94.0%, lowering the analysis cost from \$86.25 to \$5.19 on 2,168 repositories. In a large-scale study of 120K GitHub repositories containing over 7.3 million files, MalTotal discovered 564 previously unknown malicious repositories across multiple languages at a total cost of \$338. These results demonstrate the effectiveness, scalability, and cost-efficiency of MalTotal in mitigating large-scale code poisoning attacks.
CommentsAccepted by ISSTA'26
DOI:10.1145/3832228