arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

日志异常检测的协议敏感性评估:HDFS 和 BGL 上的组件成本与目标访问敏感性

Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL

Hang Xiao, Janet Sung, Zhaoyi Li, Gangzhen Qian, Chuhong Xu

arXiv 2610.05807首次发表:更新:

发表机构

Fortinet, Inc.; Google LLC; Sony Corporate of America(飞塔公司; 谷歌有限责任公司; 索尼美国公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过跨系统消融和组件剖析,揭示协议选择(分割、表示、解析)显著影响日志异常检测基准结论,并识别出支持可解释比较所需的协议字段。

AI 中文摘要

协议选择即使在检测器设置固定的情况下也能改变从日志异常检测基准中得出的结论。我们使用 Hadoop 分布式文件系统(HDFS)和 Blue Gene/L(BGL)日志上的六种固定计数、序列和语义配置,对分割构建、表示可见性和组件成本进行了联合实证研究。随机分割将几种配置置于平均精度上限附近,而组不相交的 HDFS 和按时间顺序的 BGL 评估则产生较低的分数和不同的观察排序。在固定的 BGL 截断点,解析器选择在语义 XGBoost 平均精度上跨越 0.124,同时保持其相对于计数 XGBoost 的领先优势;最早的滚动周期逆转了该排序。一项双因素跨系统消融研究将仅源表示与通过表示语料库和逆文档频率对未标记目标模板的离线转导访问进行了对比:HDFS 到 BGL 的平均精度从仅源访问的 0.191 提升到联合语料库、目标 IDF 访问的 0.325,中间条件揭示了在固定审查预算下平均精度和检索中的方向依赖交互。组件级剖析将解析和表示成本与分类器训练、预测和存储分开。总之,这些发现将检测器比较与测试群体、预处理状态、可见信息和测量的流水线阶段联系起来,并确定了需要与分数一起报告的协议字段,以支持日志异常检测准确性和资源使用的可解释比较。

英文摘要

Protocol choices can change the conclusions drawn from log anomaly detection benchmarks even when detector settings are fixed. We present a joint empirical study of split construction, representation visibility, and component costs using six fixed count, sequence, and semantic configurations on Hadoop Distributed File System (HDFS) and Blue Gene/L (BGL) logs. Random splits place several configurations near the average-precision ceiling, whereas group-disjoint HDFS and chronological BGL evaluation produce lower scores and different observed orderings. At a fixed BGL cutoff, parser choice spans 0.124 in semantic XGBoost mean average precision while preserving its lead over count XGBoost; the earliest rolling period reverses that ordering. A two-factor cross-system ablation contrasts source-only representations with offline transductive access to unlabeled target templates through the representation corpus and inverse document frequency: HDFS-to-BGL mean average precision moves from 0.191 with source-only access to 0.325 with union-corpus, target-IDF access, and the intermediate conditions reveal direction-dependent interactions in average precision and retrieval at fixed review budgets. Component-level profiling separates parsing and representation costs from classifier training, prediction, and storage. Together, these findings connect detector comparisons to the test population, preprocessing state, visible information, and measured pipeline stages, and identify the protocol fields needed alongside a score to support interpretable comparisons of log anomaly detection accuracy and resource use.

CommentsAccepted at DASC 2026. 8 pages, 1 figure, 8 tables. Reproduction support artifact: https://doi.org/10.5281/zenodo.23151002

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑