发表机构
ResearchLab(78研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ALIBI攻击,通过向二进制文件注入虚假安全叙述的只读节,使LLM恶意软件分析器将恶意样本误判为良性,并验证了跨模型与格式的迁移性,强调需来源检查防御。
AI 中文摘要
大型语言模型正被集成到恶意软件分流工作流中,作为推理组件,总结静态证据并生成面向分析人员的判定结果。本文表明,相同的推理能力引入了一个新的攻击面。我们提出了ALIBI,一种针对前沿基于LLM的恶意软件分析器的语义掩护故事攻击。ALIBI向编译后的二进制文件添加一个小的、不执行的只读节,其中包含连贯但虚假的安全产品叙述,而不改变导入或可执行行为。它不向模型发出直接指令,而是将可疑证据重新框架为良性端点安全工具的预期行为。在包含50个恶意样本的固定PE集上,该载荷使Gemini 2.5 Pro上35个基线恶意样本中的30个翻转为良性,而GPT-5.5 Pro和Claude Opus 4.7产生显著的严重性降级,即使在判定标签被保留时,置信度也显著降低。该攻击可迁移到ELF二进制文件,其中Gemini翻转了40个中的16个。一个验证引导的防御提示大致将良性判定减半,但42.9%的恶意样本仍达到良性。因此,LLM恶意软件分析器需要来源检查,将已验证事实与攻击者控制的声明分开,而非依赖叙述信任。
英文摘要
Large language models are being integrated into malware triage workflows as reasoning components that summarize static evidence and produce analyst-facing verdicts. This paper shows that the same reasoning capability introduces a new attack surface. We present ALIBI, a semantic cover story attack against frontier LLM-based malware analyzers. ALIBI adds a small, non-executed read-only section to a compiled binary, containing a coherent but false security product narrative, without altering imports or executable behavior. Instead of issuing direct instructions to the model, it reframes suspicious evidence as expected behavior of a benign endpoint security tool. On a frozen PE set of 50 malicious samples, the payload flips 30 of the 35 baseline-malicious samples to benign on Gemini 2.5 Pro, while GPT-5.5 Pro and Claude Opus 4.7 produce substantial severity downgrades with significant confidence reductions even when verdict labels are preserved. The attack transfers to ELF binaries, where Gemini flips 16 of 40. A verification-guided defense prompt roughly halves the benign verdicts, but 42.9 percent of malicious samples still reach benign. LLM malware analyzers therefore require provenance checks that separate verified facts from attacker-controlled claims, not narrative trust.
Comments12 pages, 4 figures, 4 tables