发表机构
Department of ICT, University of Agder, Norway(信息与通信技术系,阿格德大学,挪威)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对PDF文件易被嵌入恶意代码的问题,提出基于可解释Tsetlin机的框架,通过静态分析提取特征并基于规则学习分类,在RIT - PDFMal - 2026数据集上准确率达98.02%,兼具检测性能、效率与可解释性,是实用的检测方案。
AI 中文摘要
在数字时代,便携式文档格式(PDF)因其平台独立性和丰富功能,成为存储和交换数字文档最广泛使用的文件格式之一。然而,这些特性也使PDF文件成为网络攻击者的诱人攻击目标,他们在看似合法的文档中嵌入恶意代码以破坏目标系统。本文提出了一种基于可解释的Tsetlin机(TM)的新型框架用于PDF恶意软件检测。该框架通过静态分析从PDF文档中提取显著特征,无需执行文件,并采用基于规则的学习来准确分类良性和恶意PDF文档。在RIT - PDFMal - 2026数据集上的数值评估表明,该框架具有竞争力,与多个机器学习分类器和现有方法相比,准确率达到98.02%。此外,该框架通过透明解释其分类决策提供内在可解释性。竞争的检测性能、计算效率和内在可解释性的结合,使该框架成为实际PDF恶意软件检测的有前途的解决方案。
英文摘要
In the digital era, Portable Document Format (PDF) is one of the most widely used file formats for storing and exchanging digital documents due to its platform independence and rich functionality. However, these same capabilities have also made PDF files an attractive attack vector for cyberattackers, who embed malicious code within seemingly legitimate documents to compromise target systems. This paper presents a novel interpretable Tsetlin Machine (TM)-based framework for PDF malware detection. The proposed framework extracts salient features from PDF documents through static analysis without executing the files and employs rule-based learning to accurately classify benign and malicious PDF documents. Numerical evaluation on the RIT-PDFMal-2026 dataset demonstrates that the proposed framework achieves an accuracy of 98.02%, outperforming several state-of-the-art machine learning classifiers. Moreover, the proposed framework provides intrinsic interpretability by transparently explaining its classification decisions. Edge deployment on a Raspberry Pi further supports real-time, on-device PDF malware detection. The combination of better accuracy, computational efficiency, and intrinsic interpretability makes the proposed framework a promising solution for practical PDF malware detection.
Comments8 pages, 17 figures, 7 tables. Submitted to the IEEE Symposium Series on Computational Intelligence, 2027