发表机构
CUHK; Microsoft Research Asia; Microsoft; UIUC(香港中文大学; 微软亚洲研究院; 微软; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RoutingBench评估智能体路由路径分析的可扩展性,发现基于“探索更多;消化更少”原则的技能可使分析扩展到5万路由器网络并达到99.5%准确率,同时揭示了智能体在复杂场景下的能力边界。
AI 中文摘要
近期AI模型和智能体技术的进展使得AI用于网络运维(NetOps)成为可能。然而,在分析超大规模网络(包含数百个数据中心,每个数据中心容纳数千台网络设备)时,可扩展性仍是智能体NetOps的关键瓶颈。该可扩展性挑战源于许多NetOps任务的需求,这些任务必须对设备行为的局部变化如何影响所有相关路由路径进行全局推理,即路由路径分析。本文研究这一可扩展性问题,并评估不同的智能体方法(即上下文学习、迭代推理和智能体技能)如何将路由路径分析扩展到大型复杂网络。我们提出了RoutingBench,用于评估智能体路由路径分析,涵盖不同网络规模和复杂度以及各种类型的设备变化。我们的结果表明,智能体分析前景广阔:基于“探索更多;消化更少”原则整理的智能体技能,能够在包含5万个路由器的超大规模网络上实现99.5%的路由路径分析准确率,显著超越了传统符号分析的可扩展性。同时,RoutingBench也揭示了AI智能体在复杂数据中心间网络和复合变化上的能力边界,为AI和智能体研究提出了开放性挑战。
英文摘要
Recent advances in AI models and agentic technologies make AI for network operations (NetOps) within reach. However, scalability remains a key bottleneck of agentic NetOps when analyzing hyperscale networks, which comprise hundreds of datacenters, each housing thousands of network devices. The scalability challenge is rooted in the requirement of many NetOps tasks that must conduct global reasoning on how a local change of device behavior affects all relevant routing paths, known as routing-path analysis. This paper studies this scalability problem and evaluates how different agentic approaches, namely in-context learning, iterative reasoning, and agent skills, can scale routing-path analysis to large, complex networks. We present RoutingBench for evaluating agentic routing-path analysis, with varying network size and complexity, for various types of device changes. Our results show that agentic analysis is promising: agent skills curated with a principle termed "explore more; digest less" enables routing-path analysis on hyperscale networks of 50K routers with an accuracy of 99.5%, significantly outreaching the scalability of traditional symbolic analysis. Meanwhile, RoutingBench also reveals the boundary of AI agent capability on complex inter-datacenter networks and compound changes, posing open challenges for AI and agentic research.