AI 中文总结
研究网络安全日志大语言模型分析中的上下文污染问题(被动提示注入漏洞),提出LogInject框架评估威胁,通过实验得出攻击成功率等数据,引入上下文拼接技术,评估分层防御效果,揭示需深度防御架构和人工监督。
AI 中文摘要
大语言模型越来越多地部署在安全运营中心用于日志分析任务,如总结、警报分类和威胁调查。这些系统摄取来自面向外部服务的日志并将网络日志作为自然语言上下文处理以生成安全见解。我们证明这种架构模式引入了一个关键漏洞:对手可在日志生成字段中嵌入提示注入有效载荷,其在存储中持续存在并在分析师查询大语言模型时执行,即被动提示注入。我们提出LogInject,一个评估这些威胁的系统框架。使用包含2569个对抗样本的12847条日志条目的基准LogInject-1.0,我们针对四个攻击目标评估了三个生产大语言模型。我们的发现揭示在基线条件下高达88.2%的攻击成功率。我们引入上下文拼接技术,通过跨多个日志条目分割有效载荷来逃避无状态过滤器,同时利用大语言模型的长上下文推理,成功率达76.4%。作为缓解措施,我们评估了分层防御,虽仍有8.4%的残余漏洞,但攻击减少了90.4%。我们的结果表明基于大语言模型的日志分析存在固有混淆代理漏洞,需要深度防御架构和持续的人工监督。
英文摘要
Large Language Models are increasingly deployed in Security Operations Centers for log analysis tasks including summarization, alert triage, and threat investigation. These systems ingest logs from external-facing services and process network logs as natural language contexts to generate security insights. We demonstrate that this architectural pattern introduces a critical vulnerability: adversaries can embed prompt injection payloads in log-generating fields that persist in storage and are executed when analysts query the LLM, achieving what we term passive prompt injection. We present LogInject, a systematic framework for evaluating these threats. Using LogInject-1.0, a benchmark of 12,847 log entries including 2,569 adversarial samples, we evaluate three production LLMs across four attack objectives: activity concealment, false positive generation, information exfiltration, and output hijacking. Our findings reveal an up to 88.2% attack success rate (83.4% average across models) under the baseline conditions. We introduce Context Stitching, a novel technique that fragments payloads across multiple log entries to evade stateless filters while exploiting LLM long-context reasoning, achieving a 76.4% success rate. As mitigation, we evaluate layered defenses by combining input filtering, prompt hardening, and output validation, demonstrating a 90.4% attack reduction, although 8.4% residual vulnerability persists. Our results establish that LLM-based log analysis creates an inherent confused deputy vulnerability where untrusted data and trusted instructions compete indistinguishably for model attention, requiring defense in-depth architectures and continued human oversight for security-critical decisions.