AI 中文总结
LOGOS 利用实体-事件共现和时间前驱关系,将原始日志压缩为前驱森林,在零先验知识下实现故障传播隔离,显著降低噪声并提升根因定位准确率。
AI 中文摘要
商业可观测性平台依赖领域工件,如分布式追踪、拓扑图和基线指标。然而,在排查专有软件故障时,企业运维人员只能获得原始、无标注的文本日志。我们探索了仅凭日志诊断的极端边界:仅使用原始文本日志,我们能在多大程度上隔离故障传播?我们提出 LOGOS,一个无监督系统,利用实体-事件共现和时间前驱关系,将数百万条原始日志行压缩为紧凑的前驱森林。在 25 次生产环境企业故障和 12 个开源问题上的评估中,LOGOS 在零先验知识下运行——无需种子查询、观察到的症状或预定义的事故边界。在中位墙钟时间 4.5 分钟内,LOGOS 消除了中位 99.8% 的背景噪声,实现了 0.76 的平均召回率,并以中位 16 小时的诊断验证提前时间检测故障级联——将告警洪流压缩 124 倍,使 80% 的企业(100% 的开源)零样本 LLM 根因准确率得以实现。
英文摘要
Commercial observability platforms rely on domain artifacts like distributed traces, topology maps, and baseline metrics. However, when troubleshooting proprietary software, enterprise operators are left with only raw, unannotated text logs. We explore the extreme boundary of log-only diagnosis: To what extent can we isolate failure propagation using strictly raw text logs? We present LOGOS, an unsupervised system that exploits entity-event co-occurrence and temporal precedence to collapse millions of raw log lines into a compact precedence forest. Evaluated across 25 production enterprise outages and 12 open-source issues, LOGOS operates with zero prior knowledge---requiring no seed queries, observed symptoms, or pre-defined incident boundaries. In a median wall-time of 4.5 minutes, LOGOS eliminates a median 99.8% of background noise, achieves 0.76 mean recall, and detects failure cascades with a 16-hour median diagnosis-verified lead time---consolidating alert floods 124x to enable 80% enterprise (100% open-source) zero-shot LLM root-cause accuracy.
Comments20 pages, 3 figures, 13 tables