arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09179cs.CRcs.SE

Malaika:通过三重基础的智能推理理解恶意软件

Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning

Xingzhi Qian, Xinran Zheng, Yiling He, Lorenzo Cavallaro

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对恶意软件理解挑战,将其视为基础推理问题,提出需三种基础形式。介绍多智能体框架Malaika,通过受分析师启发的推理等实现三种基础机制,用于安卓恶意软件分析,实验表明该框架提高分析质量,为可靠恶意软件理解及软件分析提供基础。

中文摘要 AI 辅助

最近基于大语言模型(LLM)的系统在以安全为重点的代码分析方面展现出了有前景的能力。然而,恶意软件理解带来了独特挑战:分析师必须在部分可观测性下,从与良性功能交织的稀疏、分散证据中重建高级恶意行为。虽然静态分析能揭示安全相关信号,但核心挑战不仅是识别可疑代码,还要确定证据是否足以支持可审计的行为级结论。我们将恶意软件理解表述为一个基础推理问题,并认为可靠的行为重建需要三种互补的基础形式。领域基础约束行为假设的生成和评估方式,语义基础定位并连接支持程序的证据,知识基础通过外部可验证的威胁知识支持行为归因。为研究该假设,我们提出了Malaika,这是一个多智能体框架,通过受分析师启发的推理、工具介导的证据定位和基于检索的行为归因来实现这三种基础机制。我们将Malaika实例用于安卓恶意软件分析,并在恶意软件理解任务上对其进行评估。结果表明,Malaika比之前基于LLM的恶意软件分析框架提高了分析质量,并表明可靠性不仅取决于模型能力,还取决于推理过程。特别是,与恶意软件分析系统和前沿智能框架的比较表明,基础感知推理能产生更精确和可审计的结论。消融研究进一步支持了基础假设。这些发现表明,基础感知推理为可靠的恶意软件理解提供了一个有原则的基础,更广泛地说,为基于证据的软件分析提供了基础。

英文摘要

Recent LLM-based systems have shown promising capabilities for security-focused code analysis. Malware understanding, however, poses a distinct challenge: analysts must reconstruct high-level malicious behaviors under partial observability from sparse, dispersed evidence intertwined with benign functionality. While static analysis can expose security-relevant signals, the central challenge is not merely identifying suspicious code, but determining whether the evidence sufficiently supports an auditable behavior-level conclusion. We formulate malware understanding as a grounded reasoning problem and argue that reliable behavior reconstruction requires three complementary forms of grounding. Domain grounding constrains how behavior hypotheses are generated and evaluated, semantics grounding localizes and connects supporting program evidence, and knowledge grounding supports behavioral attribution through externally verifiable threat knowledge. To study this hypothesis, we present Malaika, a multi-agent framework that operationalizes the three grounding mechanisms through analyst-inspired reasoning, tool-mediated evidence localization, and retrieval-based behavioral attribution. We instantiate Malaika for Android malware analysis and evaluate it on malware-understanding tasks. Results show that Malaika improves analysis quality over prior LLM-based malware-analysis frameworks and demonstrate that reliability depends not only on model capability but also on the reasoning process. In particular, comparisons against malware-analysis systems and frontier agentic frameworks show that grounding-aware reasoning produces more precise and auditable conclusions. Ablation studies further support the grounding hypothesis. These findings suggest that grounding-aware reasoning provides a principled foundation for reliable malware understanding and, more broadly, for evidence-grounded software analysis.

↑