发表机构
Stevens Institute of Technology; Nokia Bell Labs(史蒂文斯理工学院; 诺基亚贝尔实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 ASTRA 智能体系统,通过协调三类专业智能体与优化循环生成证据支持的工单排查报告,在 987 个电信故障工单上取得 4.13/5.0 的平均质量得分,硬件故障诊断存在文本证据渠道的根本局限。
AI 中文摘要
技术运营团队通过整合工单文本、历史案例、系统日志和技术文档中的碎片化证据来处理大量事件。现有自动化方案常依赖整体生成,未进行显式证据建模或溯源,当各来源的关键信号稀疏时,输出难以验证。我们提出 ASTRA,这是一个用于工单解决的智能体系统,其中中央协调器协调三个专业信息收集智能体,并驱动法官-协调器优化循环以生成有证据支持的故障排查报告。TicketSimilarityAgent 通过密集检索和大语言模型(LLM)重排序检索相关历史先例;LogAgent 使用确定性过滤和受约束的 LLM 分析将数十万条日志行提炼为结构化、有引用依据的发现;DomainKnowledgeAgent 通过模型上下文协议(MCP)检索相关技术知识。它们的输出被转换为主张-证据表示,将每个主张与逐字来源段落关联,分配支持级别,并防止交叉归因。JudgeAgent 根据五个标准对报告评分,OrchestratorAgent 将低分转化为针对性的后续查询,以进行有限的迭代优化。在跨七个产品线的 987 个真实电信故障工单上评估,ASTRA 平均质量得分为 4.13/5.0,59.9% 的报告在组件系列级别或更高级别识别出故障区域;相关性和清晰度得分分别为 4.88 和 4.94,而错误案例中伪造技术细节占比低于 3%。按故障类型分层分析显示,硬件故障比软件或配置故障更难(科恩 d=0.80),这表明基于文本的证据渠道在硬件故障诊断中存在根本局限。
英文摘要
Technical operations teams resolve large volumes of incidents by synthesizing fragmented evidence from ticket text, historical cases, system logs, and technical documentation. Existing automation often relies on monolithic generation without explicit evidence modeling or provenance, making outputs difficult to verify when critical signals are sparse across sources. We propose ASTRA, an agentic system for ticket resolution in which a central orchestrator coordinates three specialist information-gathering agents and drives a judge-orchestrator refinement loop to produce evidence-backed troubleshooting reports. TicketSimilarityAgent retrieves relevant historical precedents through dense retrieval and LLM reranking; LogAgent distills hundreds of thousands of log lines into structured, quote-grounded findings using deterministic filtering and constrained LLM analysis; and DomainKnowledgeAgent retrieves relevant technical knowledge via the Model Context Protocol (MCP). Their outputs are transformed into a claim-evidence representation linking each claim to a verbatim source passage, assigning a support level, and preventing cross-attribution. A JudgeAgent scores the report on five criteria, while the OrchestratorAgent converts low scores into targeted follow-up queries for bounded iterative refinement. Evaluated on 987 real-world telecom fault tickets across seven product lines, ASTRA achieves a mean quality score of 4.13/5.0, with 59.9% of reports identifying the fault area at the component-family level or better. Relevance and Clarity scores are 4.88 and 4.94, respectively, while fabricated technical details remain below 3% of error cases. Stratification by fault type reveals that hardware faults remain substantially harder than software or configuration faults (Cohen's d=0.80), pointing to a fundamental limitation of text-based evidence channels for hardware fault diagnosis.