arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21306cs.LG

快速且准确的文本内容文件类型识别

Fast And Accurate Text Content File Type Identification

  • CrowdStrike, Inc.(CrowdStrike公司)
  • Univ. of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

Manu Nandan, Michael Brautbar, Edward Raff

AI总结:

本研究提出一种神经网络模型,用于快速准确地识别文本内容文件类型,在开源文件上比 Magika 准确率更高、速度快约四倍且体积小 28%。

AI中文摘要:

各组织普遍需要一种能够基于文件内容识别文件类型的工具,尤其是在网络安全领域,因为文件扩展名和魔数不可信。现有工具在实践中表现良好,但在计算负载和检测时间方面(如基于模型的工具 Magika)或检测准确性方面(如使用编程语言结构的文件解析工具)仍有很大改进空间。在本研究中,我们提出了一种用于识别文本内容文件(尤其是源代码)类型的神经网络模型,该模型比其他现有工具更准确且更快。我们在开源文件上的实验表明,该模型在文本内容文件类型识别上的平均准确率更高,同时速度约为 Magika 的四倍,而模型大小缩小了 28%。

英文摘要:

A common requirement across organizations is to have a tool that can identify file types based on their contents, particularly in the cybersecurity domain where magic numbers and file extensions can not be trusted. While existing tools work well in practice, there is plenty of room for improvement either in terms of computational load and time for detection in the case of model based tools like Magika or in terms of accuracy of detection in the case of file parsing tools that use programming language constructs. In this study, we propose a neural network model for identification of types of text content files, especially source code, that is more accurate and faster than other available tools. Our experiments on open-source files indicate that it is not only more accurate on average for text-content file-type identification, but also approximately four times faster than Magika, while being 28% smaller in size.

补充信息

↑