arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向互联网规模服务的基于受限创造力的智能体根因分析

Agentic RCA for Internet-Scale Services Using Constrained Creativity

Sayan Sinha, Vipul Harsh, B. Aditya Prakash, Vyas Sekar, Hui Zhang

arXiv 2610.08622首次发表:更新:

发表机构

Georgia Tech; Carnegie Mellon University; Conviva(佐治亚理工学院; 卡内基梅隆大学; 康维瓦)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对互联网规模服务故障排查,提出基于受限创造力范式的E4智能体系统,结合LLM自动化与结构化DSL,实现高准确率、可解释且低成本,准确率提升达62%,成本降低12倍。

AI 中文摘要

互联网规模服务的系统管理员需要解决故障事件以维持此类服务的可靠性。理想情况下,我们希望一个故障排查系统能够:(1) 对已知和未知事件具有高准确率的表达能力;(2) 在规模上具有成本效益;(3) 可解释,以提供操作员可采取行动的可行见解;(4) 对操作员的工作量要求低。不幸的是,大多数现有系统,包括新兴的LLM辅助智能体工作流和用于编写多样化RCA算法的结构化框架,都未能同时满足这四个要求。我们提出了E4,一个用于互联网规模服务故障排查的新型智能体系统。E4体现了受限创造力的范式,将LLM辅助自动化与探索的最佳能力同结构化方法的可解释性和效率相结合。我们不是允许LLM智能体编写任意代码或生成任意响应,而是为智能体提供一个受限的DSL,通过简单的无循环数据流程序生成其响应。该DSL配备了用于故障排查的高级操作符,使E4的输出准确、可验证且可解释。在合成和真实工作负载的混合测试中,与最先进的解决方案相比,E4实现了高达62%的准确率提升,同时以高达12倍的成本降低提供了更具可解释性的响应。

英文摘要

System administrators of Internet-scale services need to resolve failure incidents to maintain reliability of such services. Ideally, we want a troubleshooting system to be: (1) expressive to known and unknown incidents with high accuracy; (2) cost efficient at scale; (3) explainable to provide actionable insights operators can act on; and (4) entail low effort from the operators. Unfortunately, most existing systems, including emerging LLM-assisted agentic workflows and structured frameworks for authoring diverse RCA algorithms fall short of achieving all four requirements. We present E4, a novel agentic system for troubleshooting for Internet-scale services. E4 embodies the paradigm of constrained creativity that combines the best of LLM-assisted automation and exploration with the explainability and efficiency of a structured approach. Instead of allowing an LLM agent to write arbitrary code or generate arbitrary responses, we provide the agent a restricted DSL to generate its response via simple loop-free data flow programs. This DSL, equipped with high level operators for troubleshooting, makes E4's output accurate, verifiable and explainable. On a mix of synthetic and real-world workloads, E4 achieves up to 62% better accuracy compared to state-of-the-art solutions, while providing more explainable responses at up to 12x reduced cost.

Comments21 pages, including the references and appendix; 10 figures; 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑