WirelessOpsAgent:无线网络行动保障的基准测试与智能体设计
WirelessOpsAgent: A Benchmark and Agent Design for Action Assurance in Wireless Networks
AI总结:
针对现有无线网络运维LLM智能体基准未测试执行阶段支持检查的问题,本文提出WirelessOptBench基准与WirelessOpsAgent智能体,使后者在600片段评估中精确行动准确率达0.983,不安全执行率大幅下降。
AI中文摘要:
大型语言模型(LLM)智能体正逐渐成为自主无线网络运维的规划工具。然而,若支持的遥测数据过时或不一致,即使在提案时正确的任务答案,在执行阶段仍可能存在安全隐患。现有基准主要基于固定观测评估任务解决能力,未对执行阶段的支持检查进行测试。本文提出WirelessOptBench,这是一个用于无线网络行动保障的基准测试,它将无线任务转化为带有可控遥测故障和行动约束的执行状态决策片段。我们还开发了WirelessOpsAgent,该智能体基于当前证据确定候选行动,并在执行前修复可恢复的支持故障。在使用600个片段的三个骨干评估中,WirelessOpsAgent的精确行动准确率最高达0.983;相较于最安全的基线,在Claude Sonnet 4.6上,不安全执行率从82.2%降至10.3%。我们将WirelessOptBench公开于此https URL。
英文摘要:
Large language model (LLM) agents are emerging as planners for autonomous wireless network operations. Yet a task answer that is correct at proposal time can still be unsafe at execution time if supporting telemetry is stale or inconsistent. Existing benchmarks mainly evaluate task solving from fixed observations and leave support checking at execution time untested. We introduce WirelessOptBench, a benchmark for action assurance in wireless operations. It turns wireless tasks into execution state decision episodes with controlled telemetry faults and action constraints. We further develop WirelessOpsAgent, which grounds candidate actions in current evidence and repairs recoverable support failures before execution. Across three backbone evaluations with 600 episodes each, WirelessOpsAgent achieves up to 0.983 Exact Action Accuracy. On Claude Sonnet 4.6, the Unsafe APPLY Rate decreases from 82.2% to 10.3% relative to the safest baseline. We make WirelessOptBench available at https://anonymous.4open.science/r/wirelessopsbench-artifact-D969/.