发表机构
Huawei Technologies; Paris Research Center, Huawei Technologies; Queen Mary University of London(华为技术有限公司; 华为技术有限公司巴黎研究中心; 伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出CTBench基准,评估AI智能体的电信故障排查能力,发现其路径恢复表现好于根本原因分析,且资源消耗与诊断效果无必然关联。
AI 中文摘要
智能体正日益被考虑用于自动化网络运维与维护,在此场景中,工程师必须诊断网络故障、优化配置以提升服务质量、降低运维成本,同时需在严格约束下开展工作。然而,现有评估无法准确建模真实网络特性,也无法在具备不同厂商、设备、协议及接口的部分可观测电信环境下评估智能体。本文中,我们提出CTBench,这是一个用于评估智能体是否具备合格电信故障排查工程师能力的公开基准。CTBench聚焦于根本原因分析与路径恢复,每项任务均由专家构建并标注了包含黄金证据步骤的丰富任务元数据,其采用基于专家实践的指标,同时评估最终答案与诊断证据。对代表性框架-模型组合的实验表明,最先进的智能体在路径恢复任务中识别端点的表现极佳,但在根本原因分析方面整体表现欠佳,尤其在接口状态、链路层、服务管理及其他运维故障上存在困难。最重要的是,即便智能体给出合理或正确的最终答案,也常无法提供运维实践所需的基于证据的诊断。我们的结果进一步显示,路径恢复通常资源消耗更高,但资源使用量更大并不必然带来更好的诊断效果。
英文摘要
Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public benchmark for assessing whether an agent behaves like a competent telecom troubleshooting engineer. CTBench focuses on root cause analysis and path restoration. Each task is constructed by experts and annotated with rich task metadata, including golden evidence steps. CTBench uses expert-grounded metrics that evaluate both final answers and the diagnostic evidence. Experiments with representative harness-model combinations show that state-of-the-art agents perform very well at identifying endpoints in path-restoration tasks but, more generally, underperform in root cause analysis. In particular, agents struggle with interface state, link-layer, service-management, and other operational faults. Most importantly, even when agents produce plausible or correct final answers, they often fail to provide the evidence-grounded diagnoses required in operational practice. Our results further show that path restoration is generally more resource expensive, yet larger resource usage does not necessarily translate into better diagnosis.