arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23179cs.NIcs.AI

NetConfArena:面向闭环网络配置的LLM智能体可执行基准

NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Chang Liu, Xiaohui Xie, Xinyi Chen, Yong Cui

AI总结:

本文提出NetConfArena基准,在仿真多设备网络中评估LLM智能体的闭环网络配置能力,发现其失败不限于命令错误,还存在规范遵守等缺陷,为相关研究指明改进方向。

AI中文摘要:

大型语言模型(LLM)智能体在网络配置自动化领域的吸引力日益提升,但其可靠性与失败模式却鲜为人知。评估此类智能体的必要前提是在真实且无风险的环境中开展测试,然而现有基准存在不足:它们常将网络配置视为静态命令生成,或依赖过于简化的设置,这类评估低估了网络配置的核心挑战——其正确性需要对协议复杂性与拓扑依赖性进行推理。本文提出NetConfArena,一款用于评估闭环网络配置中LLM智能体的可执行基准。NetConfArena将智能体置于仿真多设备网络中,提供标准化且紧凑的任务执行动作接口,并通过隐藏的任务特定可执行测试用例评估网络行为。该基准依赖LLM辅助、仿真支撑的流水线,将面向人类的网络材料转换为可复用的参数化任务模板。我们在96个协议聚焦型任务模板实例化的480个任务实例上评估代表性LLM智能体,共生成3840条执行轨迹,结果表明失败不仅限于命令错误,还暴露了智能体在任务规范遵守、鲁棒规划与执行方面的缺陷。这些发现指明两个未来方向:一是使用经过验证的轨迹作为监督信号改进基础模型,二是设计使智能体执行更可靠、可问责的管控机制。

英文摘要:

Large language model (LLM) agents are increasingly attractive for automating network configuration, yet their reliability and failure patterns are poorly understood. An essential prerequisite is to assess such agents in a realistic but risk-free environment. Existing benchmarks, however, fall short: they often treat configuration as static command generation or rely on overly simplified settings. Such evaluations understate the core challenges of network configuration, where correctness requires reasoning about protocol complexity and topology dependence. We present NetConfArena, an executable benchmark for evaluating LLM agents in closed-loop network configuration. NetConfArena places agents in emulated multi-device networks, provides a standardized and compact action interface for task execution, and evaluates the resulting network behavior with hidden task-specific executable test cases. The benchmark relies on an LLM-assisted, emulation-grounded pipeline, which converts human-oriented network materials into reusable parameterized task templates. We evaluate representative LLM agents on 480 task instances instantiated from 96 protocol-focused task templates, yielding 3840 execution trajectories, and show that failures are not limited to command errors. The failures also reveal gaps in task-specification adherence and robust planning and execution. These findings suggest two future directions: using validated trajectories as supervision signals to improve foundation models, and designing harness mechanisms that make agent execution more reliable and accountable.

↑