arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一跳的成本:NLIP 与 A2A 的基准测试

The Cost of a Hop: Benchmarking NLIP and A2A

Ranjan Sinha, Anindita Das, Ashika Anand Babu, Hari Palleti

arXiv 2610.04053首次发表:更新:

发表机构

IBM Software; IBM; OpenShift Networking Red Hat; Red Hat(IBM软件; IBM; OpenShift网络红帽; 红帽)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次基准测试 NLIP 与 A2A 协议,分解延迟阶段,发现 NLIP 在轻量协调中快 4-9.6 倍,优势源于连接建立,并给出协议选择指南。

AI 中文摘要

基于大型语言模型(LLM)构建的自主智能体需要标准化协议以跨系统互操作。目前已有多种协议(A2A、MCP、ACP、ANP、NLIP),但自然语言交互协议(NLIP)尚未出现在任何受控性能研究中,也没有工作测量过智能体协议的延迟消耗在何处。我们在三个独立的硬件环境中,将 NLIP 与智能体到智能体(A2A)协议进行实证比较,将延迟分解为消息创建、连接和发送三个阶段。对于轻量级协调,NLIP 在两个环境中比基线 A2A SDK 实现快 8.4-9.6 倍,在第三个环境中约快 4 倍;优势方向一致,但幅度取决于硬件。该优势具有阶段性:在端到端流水线中,LLM 推理占主导,两者接近持平。差异几乎完全来自连接建立。为了测试 A2A 的最佳表现,我们还启用了连接缓存的 A2A SDK;缓存将差距缩小了硬件相关的量,从一台机器上的 2.75 倍到更快硬件上的接近持平,在规模下缓存优化的 A2A-SDK 与 NLIP 相当。我们报告这些测量条件,而不对残余发送阶段成本给出单一因果解释。与更优化的 Python-A2A 相比,NLIP 在同一阶段领先约 4 倍。最后,我们提供一份针对工作负载特征的关键协议选择指南。

英文摘要

Autonomous agents built on Large Language Models (LLMs) need standardized protocols to interoperate across systems. Several now exist (A2A, MCP, ACP, ANP, NLIP), but the Natural Language Interaction Protocol (NLIP) has not appeared in any controlled performance study, and no work has measured where an agent protocol's latency is spent. We compare NLIP and the Agent-to-Agent (A2A) protocol empirically, decomposing latency into message creation, connection, and send phases across three independent hardware environments. For lightweight coordination, NLIP is 8.4-9.6x faster than the baseline A2A SDK implementation on two environments and about 4x on a third; the direction of the advantage is consistent, its magnitude depends on the hardware. The advantage is stage-specific: for the end-to-end pipeline, where LLM inference dominates, the protocols are near parity. The difference comes almost entirely from connection setup. To test A2A at its best, we also ran A2A SDK with connection caching enabled; caching narrows its gap with NLIP by a hardware-dependent amount, from 2.75x on one machine to near-parity on faster hardware, where at scale a cache-optimized A2A-SDK matches NLIP. We report these as measured conditions without a single causal account of the residual send-phase cost. Against the more optimized Python-A2A, NLIP leads by about 4x on the same stage. We close with a protocol-selection guide keyed to workload characteristics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑