arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14434cs.CRcs.SE

恶意软件分析沙盒中人工智能环境现实差距的测量研究

A Measurement Study of AI-Environment Realism Gaps in Malware-Analysis Sandboxes

Zhiyong Sui, Lamine Noureddine, Mst Eshita Khatun, Sideeq Bello, Babangida Bappah, Justin Woodring, Aisha Ali-Gombe

首次发表
浏览论文内容

中文总结 AI 辅助

研究恶意软件分析沙盒中人工智能环境现实差距,通过AIprint框架提取工件,在多后端和参考主机上评估,发现传统基线难区分真实系统与沙盒,揭示操作不对称,即重现环境比检测欺骗更昂贵。

中文摘要 AI 辅助

沙盒技术仍然是观察可疑程序行为的核心技术,但具有环境感知能力的恶意软件在怀疑被分析时越来越多地抑制执行。前代沙盒逃避技术主要集中在虚拟化工件、时间差异和磨损逼真度上。本文首次对人工智能环境工件作为新的沙盒逃避表面进行了系统测量研究。通过AIprint探测框架实现对现实差距的操作化,该框架捕获人工智能软件生态系统留下的持久工件。从GitHub上的284个开源人工智能项目中系统提取450个独特工件,编译成无特权的Windows探测程序,并在七个商业和开源沙盒后端以及三个具有人工智能能力的参考主机上进行评估。结果表明传统的虚拟机检测基线无法可靠区分真实的人工智能系统和现代沙盒,有12个人工智能环境工件出现在参考主机上而未出现在评估的后端。控制安装实验建立了人工智能工具和包安装与可测量的人工智能环境工件积累之间的因果关系,自适应欺骗实验揭示了基本的操作不对称性:重现令人信服的人工智能软件环境比检测浅层欺骗要昂贵得多。

英文摘要

Sandboxing remains a core technique for observing suspicious program behavior, yet environment-aware malware increasingly suppresses execution when analysis is suspected. Prior generations of sandbox evasion focused on virtualization artifacts, timing discrepancies, and wear-and-tear realism. In this paper, we present the first systematic measurement study of AI-environment artifacts as a new sandbox-evasion surface. We operationalize this realism gap through AIprint, a probe framework that captures persistent artifacts left behind by AI-capable software ecosystems, including AI-assistant configuration directories, model caches, environment variables, local inference services, and package dependencies. We systematically extract 450 unique artifacts from 284 open-source AI projects on GitHub, compile them into unprivileged Windows probes, and evaluate them across seven commercial and open-source sandbox backends together with three AI-capable reference hosts. Our results show that traditional VM-detection baselines fail to reliably distinguish real AI-capable systems from modern sandboxes, whereas twelve AI-environment artifacts appear on the reference hosts and on none of the evaluated backends. A controlled 214-step installation experiment establishes a causal relationship between AI tool and package installation and measurable AI-environment artifact accumulation, while adaptive spoofing experiments reveal a fundamental operational asymmetry: reproducing convincing AI software environments is substantially more expensive than detecting shallow spoofing.

↑