arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TAILOR:面向长尾日志解析的模板保持增强方法

TAILOR: Template-Preserving Augmentation for Long-Tailed Log Parsing

Sepideh Hodaeian, Zhenhao Li, An Ran Chen

arXiv 2609.25261首次发表:更新:

发表机构

University of Alberta; University of Yourk(阿尔伯塔大学; 约克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对日志解析中稀有模板性能差的问题,提出TAILOR框架,通过模板保持增强丰富稀有日志组,提升解析准确率19%,且兼容不同大语言模型。

AI 中文摘要

日志解析对于系统日志分析至关重要,因为它通过将非结构化的日志消息转换为结构化的日志模板,支持调试、监控和异常检测等任务。然而,真实世界的日志数据集呈现出高度不平衡的长尾分布,其中少数频繁出现的模板占据主导地位,而许多稀有模板仅出现几次。这种不平衡导致评估结果过于乐观,因为频繁模板主导了基准指标,而稀有但操作上重要的事件上的较差性能在很大程度上被掩盖。在本文中,我们研究了稀有日志组(定义为实例数少于五个的日志组)的普遍性和影响。我们在广泛使用的Loghub-2.0基准上的实证研究表明,稀有日志组占所有模板的近20%,但占日志消息的比例不到0.01%。由于它们仅包含少量实例,所有被评估的解析器在这些组上都经历了显著的性能下降。为了解决这一挑战,我们提出了TAILOR,一个通过模板保持增强来改进稀有日志组模板推断的日志解析框架。TAILOR在模板推断之前,用与模板一致的日志消息丰富稀有日志组。额外的结构证据有助于区分静态标记和动态变量。我们的实验结果表明,TAILOR在稀有日志组上的解析准确率比最强基线提高了19%,同时在完整数据集上保持了有竞争力的性能。我们进一步表明,所提出的增强策略在不同的大语言模型骨干上具有泛化能力,并且在不修改其核心架构的情况下,持续改进现有的基于大语言模型的解析器。

英文摘要

Log parsing is essential for system log analysis because it supports tasks such as debugging, monitoring, and anomaly detection by transforming unstructured log messages into structured log templates. However, real-world log datasets exhibit highly imbalanced, long-tailed distributions, where a small number of frequent templates dominate while many rare templates appear only a few times. This imbalance causes evaluation results to be overly optimistic by allowing frequent templates to dominate benchmark metrics, while poor performance on rare yet operationally important events remains largely hidden. In this paper, we investigate the prevalence and impact of rare log groups, defined as log groups with fewer than five instances. Our empirical study on the widely used Loghub-2.0 benchmark shows that rare log groups account for nearly 20% of all templates but less than 0.01% of log messages. Because they contain only a handful of instances, all evaluated parsers experience substantial performance degradation on these groups. To address this challenge, we propose TAILOR, a log parsing framework that improves template inference for rare log groups through template-preserving augmentation. TAILOR enriches rare log groups with template- consistent log messages before template inference. The additional structural evidence helps distinguish static tokens from dynamic variables. Our experimental results show that TAILOR improves parsing accuracy on rare log groups by 19% over the strongest baseline while maintaining competitive performance on complete datasets. We further show that the proposed augmentation strategy generalizes across different LLM backbones and consistently improves existing LLM-based parsers without modifying their core architectures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑