发表机构
University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对微服务RCA的现有方法存在局限,提出图增强LLM智能体框架GALA+,结合STRIX与SURE-Score,在基准测试中性能优于最佳LLM基线,获SRE专家认可。
AI 中文摘要
微服务根本原因分析(RCA)需要关联复杂服务依赖图内异构遥测数据中的故障。现有方法常依赖单一遥测模态;近期基于大语言模型(LLM)的方法易出现无约束探索和幻觉问题;且多数系统仅停留在故障排序,无法生成可执行的事件响应。我们提出GALA+,这是一个以图引导调查为核心的图增强LLM智能体框架,利用服务依赖关系限制探索范围,并通过局部多模态证据优化诊断。对于初始假设生成,GALA+将互补遥测信号与STRIX(一种新型的感知追踪和图结构的评分模块)相结合。GALA+随后生成排序后的诊断结果、事件摘要和分层行动建议。我们还引入了SURE-Score,这是一个与行业站点可靠性工程(SRE)专家联合开发的人工引导评估框架,用于评估RCA特定输出质量,超越传统文本相似度指标。在两个微服务基准测试中,GALA+始终取得最强的整体结果,在AC@1指标上超过最佳LLM基线25个百分点以上,同时在SURE-Score和独立人工SRE评估中均获得最高评分。
英文摘要
Microservice root cause analysis (RCA) requires correlating failures across heterogeneous telemetry within complex service dependency graphs. Existing methods often rely on a single telemetry modality; recent LLM-based approaches can suffer from unconstrained exploration and hallucination; and most systems stop at fault ranking without producing actionable incident response. We present GALA+, a graph-augmented LLM agentic framework centered on graph-guided investigation, which uses service dependencies to bound exploration and refine diagnosis through localized multi-modal evidence. For initial hypothesis generation, GALA+ combines complementary telemetry signals with STRIX, a novel trace- and graph-structure-aware scoring module. GALA+ then produces ranked diagnoses, incident summaries, and stratified action recommendations. We further introduce SURE-Score, a human-guided evaluation framework co-developed with industry SRE experts for assessing RCA-specific output quality beyond conventional text similarity metrics. On two microservice benchmarks, GALA+ consistently achieves the strongest overall results, surpassing the best LLM-based baseline by more than 25 percentage points in AC@1, while also receiving the highest ratings from both SURE-Score and independent human SRE evaluation.
CommentsProceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE '26), October 12--16, 2026, Munich, Germany