AI 中文总结
研究生产微服务故障根因分析难题,提出结构化多智能体RCA管道,性能超现有方法。引入反向推理智能体和自动规则挖掘管道,发现故障证据多,瓶颈在智能体推理能力,模型推理和领域知识是主要限制,进步需模型改进。
AI 中文摘要
在生产微服务故障中识别根本原因需要对大规模、多模态遥测数据(包括指标、日志和追踪)进行推理,这一问题对传统方法和基于大语言模型(LLM)的方法都具有挑战性。OpenRCA数据集就体现了这些挑战,现有方法在该数据集上准确率一直很低。本文表明传统因果发现方法和现有基于LLM的多智能体系统在此基准测试中无法可靠识别根本原因,并提出了结构化多智能体RCA管道,其性能大幅超越现有基线,支持领域知识和无知识操作模式。为诊断故障起源,引入反向推理智能体,将故障分类为推理差距或数据模糊性。分析表明多数故障存在所需证据,瓶颈在于智能体正确推理的能力。还引入自动规则挖掘管道,减少对人工知识整理的依赖。在所有配置中,模型推理能力和领域知识是主要限制因素,即使证据提取完美,推理性能仍有实际限制,进步需要在模型层面改进。
英文摘要
Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, logs, and traces, a problem that has proved resistant to both classical and LLM-based approaches. The OpenRCA dataset exemplifies these challenges: it is large-scale, multimodal, and lacks detailed domain knowledge, and yields consistently low accuracy across all existing methods. We show that classical causal discovery methods and existing LLM-based multi-agent systems fail to reliably identify root causes on this benchmark, and present a Structured Multi-Agent RCA pipeline that substantially outperforms existing LLM-based and classical baselines, supporting both domain-knowledge and knowledge-free operating modes. To diagnose where failures originate, we introduce a reverse reasoning agent that, given the correct answer, identifies which signals in the extracted anomalies support it and determines whether Stage~1 had access to those signals, classifying each failure as Reasoning Gap (evidence present but unused) or Data Ambiguity (evidence genuinely absent). This analysis reveals that the required evidence is present in the vast majority of failures: the bottleneck is not data access but the agent's ability to reason over it correctly. We further introduce an automated rule mining pipeline that systematically extracts discrimination rules from reverse reasoning reports, reducing reliance on manual knowledge curation. Across all configurations, model reasoning capability and domain knowledge are the primary constraints: stronger models embed more domain expertise, and explicit knowledge injection partially compensates for this gap. Reasoning performance remains practically bounded even when evidence extraction is perfect: scaffold engineering and better data pipelines alone cannot close this gap; progress requires improvements at the model level.