发表机构
Mitsubishi Electric Corporation; Mitsubishi Electric Research Laboratories(三菱电机公司; 三菱电机研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对APT检测难题,提出CAPTAIN方法,利用通用预训练语言模型,经最少领域无关预处理,通过编码历史、注入上下文令牌及应用平滑滤波器,实现稳健日志评分,在基准测试中表现良好,降低了日志预处理成本。
AI 中文摘要
高级持续性威胁(APT)难以检测,因为大规模日志中只有一小部分事件与攻击相关,且调查成本高、难以扩展。先前的机器学习方法虽能减轻分析师工作量,但依赖大量精心整理的训练数据和复杂的预处理管道,构建和维护成本高。受强大APT检测基线研究启发,我们提出CAPTAIN,一种基于困惑度的探测器,利用通用预训练语言模型,经过最少的领域无关预处理,对长的、处理最少的日志条目进行稳健评分。它通过编码器模型和Q-Former风格的桥接对近期历史进行编码,将紧凑的上下文令牌注入解码器输入,使困惑度反映时间上下文。为提高稳定性,还对困惑度时间序列应用平滑滤波器。在面向APT的基准测试中,CAPTAIN与现有强大基线竞争,在输入整理较少的情况下仍保持稳健,降低了高级日志预处理的开发和运营成本。
英文摘要
Advanced Persistent Threats (APTs) remain difficult to detect because only a small fraction of events in large-scale logs are attack-related, and investigation is expensive and hard to scale. Prior machine-learning approaches can reduce analyst workload, but they often rely on heavily curated training data and sophisticated preprocessing pipelines. Building and maintaining such pipelines require substantial domain expertise and engineering cost. Motivated by insights from a study of a strong APT detection baseline, we propose CAPTAIN (Context-Augmented Perplexity-based Threat Activity log detectIoN), a perplexity-based detector that leverages general, pre-trained language models with minimal, domain-agnostic preprocessing, enabling robust scoring of long, minimally processed log entries. CAPTAIN encodes recent history with an encoder model and a Q-Former-style bridge, then injects the compact context tokens into the decoder input so that perplexity reflects temporal context. To improve stability, CAPTAIN additionally applies smoothing filters to the perplexity time series. Across APT-oriented benchmarks, CAPTAIN competes with strong existing baselines and remains robust under substantially less curated inputs, that reduces the development and operational cost of advanced log preprocessing.
Comments20 pages