arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ORCAGen:结合RAG引导的生成式AI编排上下文感知的恶意软件欺骗策略

ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI

Shihab Ahmed, Md Sajidul Islam Sajid, Teryl Taylor, Frederico Araujo, Tariqul Islam

arXiv 2610.12415首次发表:更新:

发表机构

Towson University; IBM Research; University of Maryland Baltimore County(托森大学; IBM研究院; 马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ORCAGen结合RAG与结构化提示生成恶意软件欺骗剧本,经多模型评估验证其可安全构建针对性欺骗策略,运行时轻量且确定性强,为恶意软件防御提供新途径。

AI 中文摘要

恶意软件防御通常会尽快清除或隔离可疑程序,这种方法虽能有效遏制威胁,但也会错失观察攻击者行为、部署针对性对策的机会。ORCAGen采用不同思路:它利用生成式AI离线构建恶意软件专用欺骗剧本,部署前对剧本进行验证,运行时仅执行经确认的逻辑。ORCAGen将检索增强生成(RAG)与结构化提示工程相结合,生成概念验证(PoC)恶意软件及对应的欺骗编排代码。其生成过程基于精心构建的知识库(KB),该知识库包含恶意软件流程与主动防御策略,可帮助大型语言模型(LLM)生成针对特定威胁且可执行的欺骗逻辑,而非通用或存在幻觉的输出。生成的PoC恶意软件提供了一种安全且可复现的方式,用于在将欺骗策略加入运行时剧本前,测试该策略能否破坏、重定向或抑制目标恶意软件行为。我们针对GPT-4o、GPT-5.5、Gemini 3.5 Flash、Qwen3-Coder和Claude Sonnet 4.5,从执行成功率、幻觉率、优化工作量、欺骗有效性及运行时开销等维度对ORCAGen进行评估,评估涵盖合成恶意软件场景及150个真实世界恶意软件样本,包括键盘记录器、信息窃取器和勒索软件。在所有评估场景中,GPT-5.5所需优化工作量最少,未观察到幻觉API,而Gemini 3.5 Flash实现了最低的响应时间和运行时开销。结果表明,RAG引导的结构化提示可支持规模化构建恶意软件专用欺骗剧本,这类剧本在运行时执行阶段保持轻量且具有确定性。

英文摘要

Malware defenses often remove or isolate suspicious programs as quickly as possible. While effective for containment, this approach can also waste an opportunity to observe attacker behavior and deploy targeted countermeasures. ORCAGen takes a different approach: it uses GenAI to build malware-specific deception playbooks offline, validates them before deployment, and enforces only the verified logic at runtime. ORCAGen combines Retrieval-Augmented Generation (RAG) with structured prompt engineering to generate both proof-of-concept (PoC) malware and corresponding deception orchestration code. A curated knowledge base (KB) of malware procedures and active defense strategies grounds the generation process, helping the LLM produce threat-specific and executable deception logic rather than generic or hallucinated outputs. The generated PoC malware provides a safe and reproducible way to test whether a deception strategy can disrupt, redirect, or suppress targeted malware behavior before the strategy is added to the runtime playbook. We evaluate ORCAGen across GPT-4o, GPT-5.5, Gemini 3.5 Flash, Qwen3-Coder, and Claude Sonnet 4.5 using execution success, hallucination rate, refinement effort, deception effectiveness, and runtime overhead. The evaluation covers synthesized malware scenarios and 150 real-world malware samples across keyloggers, information stealers, and ransomware. Across the evaluated scenarios, GPT-5.5 required the fewest refinements and produced no observed hallucinated APIs, while Gemini 3.5 Flash achieved the lowest response time and runtime overhead. The results show that RAG-guided structured prompting can support the scalable construction of malware-specific deception playbooks that remain lightweight and deterministic during runtime enforcement.

CommentsAccepted at 2026 IEEE 8th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS-ISA)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑