TSG Suggester:面向云事件管理的故障排除指南推荐中的树状结构知识图谱检索
TSG Suggester: Tree-Structured Knowledge-Graph Retrieval for Troubleshooting Guide Recommendation in Cloud Incident Management
浏览论文内容
中文总结 AI 辅助
针对云事件管理中TSG检索效率低的问题,提出Tree+KG方法,通过树状结构保留文档层次并融合知识图谱,在314个真实事件上实现Top 1准确率54.78%,显著优于纯文本RAG。
中文摘要 AI 辅助
大规模云服务中的值班工程师在巨大的时间压力下工作,然而为事件定位正确的故障排除指南(TSG)在很大程度上仍是一个手动、基于关键词的过程,且先前的实证研究发现,指南搜索消耗了总缓解时间的相当大一部分。我们提出了TSG Suggester,一个直接从事件描述中推荐相关TSG的检索系统。我们在一个生产事件管理系统上,针对来自18个服务团队的314个真实世界事件(涵盖112个独特的TSG),评估了五种检索策略:纯文本RAG、图像增强RAG、RAPTOR、树状结构检索,以及我们提出的树+知识图谱(Tree+KG)方法。Tree+KG将每个指南转换为保留其原生章节层次的树,将LLM生成的问题抽象附加到内部节点,以弥合指南的解决方案导向语言与事件的问题导向语言之间的差距,提取每个指南的实体知识图谱,并在查询时融合嵌入相似性与实体级匹配。Tree+KG达到了54.78%的Top 1准确率和82.48%的Top 5准确率,在每个截断点上都领先于所有基线,相比纯文本RAG在Top 1上提升了8.58个百分点。有两个发现具有独立意义。首先,结构对齐占主导地位:在需要精确区分的场景中,保留或重建文档结构的方法优于扁平分块。其次,与我们的初始假设相反,多模态增强反而有害。为指南截图添加标题并将这些标题注入,相比纯文本基线损失了22.64个Top 5点,因为通用标题稀释了嵌入而非锐化嵌入。我们报告了这两个结果的错误分析,并提供了具体的部署指导。
英文摘要
On call engineers in large scale cloud services work under intense time pressure, yet locating the correct Troubleshooting Guide (TSG) for an incident remains a largely manual, keyword driven process, and prior empirical work finds that guide search consumes a substantial fraction of total mitigation time. We present TSG Suggester, a retrieval system that recommends relevant TSGs directly from an incident description. We evaluate five retrieval strategies: Text Only RAG, Image Augmented RAG, RAPTOR, Tree Structured Retrieval, and our proposed Tree + Knowledge Graph (Tree+KG), on 314 real world incidents spanning 112 unique TSGs drawn from 18 service teams on a production incident management system. Tree+KG converts each guide into a tree that preserves its native section hierarchy, attaches LLM generated problem abstractions to internal nodes to bridge the solution oriented language of guides and the problem oriented language of incidents, extracts a per guide entity knowledge graph, and fuses embedding similarity with entity level matching at query time. Tree+KG attains 54.78% Top 1 and 82.48% Top 5 accuracy, leading every baseline at every cutoff, with an 8.58 point Top 1 gain over text only RAG. Two findings are of independent interest. First, structural alignment dominates: methods that preserve or rebuild document structure outperform flat chunking where precise discrimination matters. Second, and contrary to our initial hypothesis, multimodal enrichment actively hurts. Captioning guide screenshots and injecting the captions costs 22.64 Top 5 points relative to the text only baseline because generic captions dilute embeddings rather than sharpen them. We report error analyses for both results and provide concrete deployment guidance.
发表机构
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。