发表机构
Nankai University; Huawei Technologies Co., Ltd.; Tsinghua University(南开大学; 华为技术有限公司; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控实验比较了不同工具配置对LLM智能体在微服务根因分析中的影响,发现工具层级能提升故障类型识别但可能降低定位准确率,且收益因模型和系统而异,为工具选择提供了实证依据。
AI 中文摘要
大型语言模型(LLM)智能体正越来越多地被用于微服务系统中的根因分析(RCA),但关于如何设计和组合其工具的实证指导仍然有限。我们开展了一项受控实证研究,考察跨模型和微服务环境的工具抽象与组合。我们实现了24个结构化工具,用于指标访问(L1)、证据分析(L2)和诊断(L3),以及一个基于Python的参考设置(L0)。利用来自三个微服务系统的375个故障案例,我们评估了使用Qwen3.7-Plus的八种配置,并在四个配置的共享子集上比较了四个Qwen模型。我们的结果表明,工具配置对根因定位和故障类型识别的影响不同。使用Qwen3.7-Plus时,L3实现了82.8%的top-1定位准确率,而L0为85.4%,但每个案例所需时间不到一半。添加工具层级可以提高类型识别,但会降低定位准确率。轨迹分析揭示了智能体在咨询额外证据后覆盖正确诊断建议的案例。模型选择也会改变工具收益:将L3添加到L1+L2改善了对三个模型的诊断,但对Qwen3-8B有害,后者很少调用L3。工具配置排名还随微服务系统的不同而变化。这些发现为设计和利用RCA工具提供了实证基础,可根据模型、诊断目标和目标系统指导工具选择和组合。
英文摘要
Large language model (LLM) agents are increasingly explored for root cause analysis (RCA) in microservice systems, yet empirical guidance on how to design and combine their tools remains limited. We conduct a controlled empirical study of tool abstraction and composition across models and microservice environments. We implement 24 structured tools for metric access (L1), evidence analysis (L2), and diagnosis (L3), alongside a Python-based reference setting (L0). Using 375 failure cases from three microservice systems, we evaluate eight configurations with Qwen3.7-Plus and compare four Qwen models on a shared subset of four configurations. Our results show that tool configurations affect root cause localization and failure type identification differently. With Qwen3.7-Plus, L3 achieves 82.8% top-1 localization accuracy compared with 85.4% for L0, while requiring less than half the time per case. Adding tool levels can improve type identification while reducing localization accuracy. Trajectory analysis reveals cases in which agents override correct diagnostic recommendations after consulting additional evidence. Model choice also changes tool benefits: adding L3 to L1+L2 improves diagnosis for three models but hurts Qwen3-8B, which rarely invokes L3. Tool configuration rankings further change across microservice systems. These findings provide an empirical foundation for designing and using RCA tools, guiding tool selection and composition according to the model, diagnostic objective, and target system.