arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22723cs.NI

ANI-Gamut:跨越智能体-网络接口抽象全谱系的智能体可靠性基准测试

ANI-Gamut: Benchmarking Agent Reliability across the Gamut of Agent-Network Interface Abstractions

Lorenzo Bracciale, Pierpaolo Loreti, Andrea Mayer, Stefano Salsano, Wim Henderickx

首次发表
浏览论文内容

中文总结 AI 辅助

ANI-Gamut通过将接口抽象级别作为实验变量,在棕地基底上系统评估LLM智能体可靠性,发现类型化事务接口在实时变更任务中显著提升可靠性并降低成本,而基底隔离有效控制附带损害。

中文摘要 AI 辅助

大型语言模型(LLM)智能体日益被信任用于操作实时网络:它们读取状态、更改配置并验证结果。一个首要问题却被隐式搁置:智能体应在何种抽象级别上操作?我们将接口抽象级别设为显式、受控的实验变量,将智能体-网络接口组织成一个谱系,从原始CLI(A0)经受限包装器(A1)和标准化模型驱动配置(A2),到类型化事务性服务意图(A3-T)、对账的单一事实来源自动化(A3-R)及其组合(A4)。我们提出ANI-Gamut,一个可复现的开源实验平台,在单一、高密度部署的棕地基底上以多个级别呈现相同任务,其中许多共存服务共享资源,附带损害实际发生。我们实例化并测量了谱系中的四个点,其余仅作概念性描述,并在注入故障对基底施压时记录三个因变量:任务可靠性、对既有租户的附带损害以及运营成本。在小型运行次数的试点中(每级别n=20、n=6和n=4),接口级别显著改变可靠性和成本:在实时变更集任务上,原始shell智能体在所有六个种子上均失败,而类型化事务性接口在所有六个种子上均成功(配对McNemar p=0.031,基于六个不一致对),成本约低一个数量级,因为工程工作从智能体迁移到可复用的事务层。相比之下,附带损害在每个级别均不存在,无论良性或故障情况下:在租户隔离的基底上,智能体故障安全,爆炸半径由基底的隔离性控制,这使安全问题从智能体转移到基底。

英文摘要

Large Language Model (LLM) agents are increasingly trusted to operate live networks: they read state, change configuration, and verify the result. A first-order question is left implicit: at which level of abstraction should the agent operate? We make the interface-abstraction level an explicit, controlled experimental variable, organizing agent-network interfaces into a spectrum from raw CLI (A0) through bounded wrappers (A1) and standardized model-driven configuration (A2) to typed transactional service intent (A3-T), reconciled source-of-truth automation (A3-R), and their combination (A4). We present ANI-Gamut, a reproducible, open-source playground that exposes the same task at several levels on a single, densely populated brownfield substrate, where many coexisting services share resources and collateral damage actually arises. We instantiate and measure four points of the spectrum and describe the others only at the conceptual level, and record three dependent variables as the substrate is stressed by injected faults: task reliability, collateral damage against pre-existing tenants, and operational cost. In a pilot with small run counts (n=20, n=6 and n=4 per level), the interface level moves reliability and cost sharply: on a live change-set task a raw-shell agent fails on all six seeds while a typed transactional interface succeeds on all six (paired McNemar p=0.031, on six discordant pairs), at roughly an order of magnitude less cost, as the engineering effort migrates from the agent to a reusable transaction layer. Collateral damage, by contrast, is absent at every level, whether benign or under faults: in a tenant-isolated substrate the agents fail safe, and the blast radius is held by the substrate's isolation, which moves the safety question from the agent to the substrate.

发表机构

  • University of Rome Tor Vergata(罗马第二大学)
  • Nokia(诺基亚)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑