arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27021cs.NIcs.AI

FaulT-Bench:面向不可靠用户工单场景下网络故障诊断大语言模型智能体的基准测试

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne

首次发表
浏览论文内容

中文总结 AI 辅助

FaulT-Bench 是针对网络故障诊断 LLM 智能体的基准测试,涵盖含虚假工单的 200 个场景,评估 SADE 等智能体后发现其在错误工单上性能骤降,可用于开发可靠的智能体系统。

中文摘要 AI 辅助

基于大语言模型(LLM)的智能体被越来越多地用于网络故障诊断,但现有基准测试仅针对准确的用户工单对其进行评估,且始终假设存在故障,而这些条件在实际场景中很少能满足。我们提出 FaulT-Bench,这是一个涵盖 8 种网络拓扑的 200 个故障排查场景的基准测试,其中 5 种拓扑是从公开的从业者实验室重新实现的,覆盖真实故障、虚假故障报告、错误设备归属以及错误根本原因声明。为了隔离工单措辞对诊断的影响,我们进一步在保持网络状态固定的情况下,一次仅改变报告者的置信度和可验证细节这一因素,将 72 个错误前提工单重写为 5 种报告者角色。我们的自动化测试工具在 Kathará 中部署每个场景,让智能体通过 NIKA 工具界面进行交互,并使用 LLM 评判器从结果、修复和推理质量三个维度对自由文本诊断进行评分。通过评估 SADE、ReAct 和 Claude Code,我们发现这三者在准确工单上的表现接近饱和,且对误导具有鲁棒性,但当网络正常且工单错误时,性能会急剧下降,它们会一直探测直到将良性状况提升为根本原因,而不是得出没有问题的结论。角色重写表明,工单的书写方式比其声称的内容更重要:一份自信的错误报告的处理效果与准确报告大致相当,而一份模糊、未明确的报告会使性能急剧下降。这三个智能体的失败方式也不同,从持续过度诊断到无响应运行,成本差异很大。这些结果表明,FaulT-Bench 可作为开发能对现实世界网络故障排查中嘈杂、不可靠工单进行可靠推理的智能体系统的基准测试。

英文摘要

LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice. We present FaulT-Bench, a benchmark of 200 troubleshooting scenarios across eight network topologies, five reimplemented from public practitioner labs, spanning genuine faults, false fault reports, incorrect device attribution, and incorrect root-cause claims. To isolate how ticket wording affects diagnosis, we further rewrite 72 false-premise tickets into five reporter personas that vary reporter confidence and verifiable detail one factor at a time, holding the network state fixed. Our automated harness deploys each scenario in Kathará, lets agents interact through the NIKA tool interface, and scores free-text diagnoses with an LLM judge across outcome, fix, and reasoning quality. Evaluating SADE, ReAct, and Claude Code, we find all three are near-saturated on accurate tickets and robust to misdirection, yet degrade sharply when the network is healthy and the ticket is wrong, probing until a benign condition can be promoted to a root cause rather than concluding nothing is wrong. Persona rewrites show that how a ticket is written matters more than what it claims: a confidently wrong report is handled about as well as an accurate one, while a vague, underspecified report degrades performance sharply. The three agents also fail differently, from constant over-diagnosis to unanswered runs, at very different cost. These results position FaulT-Bench as a benchmark for developing agentic systems that can reason reliably over the noisy, unreliable tickets of real-world network troubleshooting.

发表机构

  • The University of Sydney(悉尼大学)

机构由 AI 辅助整理,请以论文原文为准。

↑