只是测试,继续前进:通过提示注入规避基于大语言模型的系统日志解释
Just Testing, Move Along: Evasion of LLM-based System Log Interpretation by Prompt Injection
浏览论文内容
中文总结 AI 辅助
研究基于大语言模型的系统日志解释中提示注入攻击问题,提出评估框架,利用真实攻击日志痕迹创建对抗示例,证明攻击可致恶意日志误分类,且LLMs解释含对抗指标可用于检测攻击。
中文摘要 AI 辅助
大语言模型(LLMs)越来越多地集成到安全运营中心(SOC)工作流程中,以支持分析师进行系统日志解释等任务。然而,LLMs直接处理不可信文本输入的能力也带来了新的攻击面。攻击者可向日志条目中注入上下文信息或明确指令,影响模型对恶意活动的解释。本文提出一个评估针对基于LLM的日志解释的提示注入攻击的框架。利用真实网络攻击中生成的日志痕迹,通过通用注入生成、细化和特定攻击优化来创建对抗性示例。评估表明这些注入可使恶意日志痕迹被误分类为良性。作为潜在补救措施,LLMs生成的解释常包含对抗性操纵指标,可用于检测此类攻击。
英文摘要
Large Language Models (LLMs) are increasingly integrated into Security Operations Center (SOC) workflows, where they support analysts in tasks such as the interpretation of system logs. However, the ability of LLMs to directly process untrusted textual input also introduces new attack surfaces. In particular, attackers can inject contextual information or explicit instructions into log entries in order to influence how malicious activity is interpreted by the model. Despite the growing adoption of LLMs for log analytics, the robustness of such systems against adversarial log injection remains largely unexplored. To address this gap, this paper presents a framework for evaluating prompt injection attacks against LLM-based log interpretation. Using log traces generated during real cyber attacks, our approach creates adversarial examples through generic injection generation, refinement, and attack-specific optimization. Our evaluation across multiple state-of-the-art LLMs shows that these injections can cause malicious log traces to be classified as benign despite containing clear indicators of compromise. As a potential remedy, we show that the explanations generated by the LLMs alongside their classifications frequently contain indicators of adversarial manipulation that can be leveraged to detect such attacks.