AI 中文总结
该研究以受控多模型试验探究MCP转A2A配置中,公开共享标签对信息流出的影响,发现公开标签与更高逐字流出相关且具强模型依赖性,相关结果非因果效应,相关工件已公开。
AI 中文摘要
针对模型上下文协议(MCP)工具使用和智能体间(A2A)委托分别评估的安全属性,无需描述单个智能体同时使用两者时的行为。我们在单一受控MCP转A2A配置中测量此类行为:测试平台驱动真实模型主机依次经过本地MCP和本地A2A环节,形成由精确确定性规则(无大语言模型评判)评分的有序事件轨迹,每次试验仅限制一个决策。在预先指定的固定三臂设计中,10种记录场景分别以“机密”标题、无标题、“公开-允许共享”标题呈现;三组设置的6个实质性记录值字节完全相同,结果为这些值中任意一个在出站消息中的逐字出现情况。4种模型×3组×4次重复共480次试验;场景为泛化单位,我们报告10个场景级值(均值、中位数、符号计数),无p值或区间。机密与无标签的对比在每个模型中均无定论且受下限限制(两组均为0或接近0),因此未显示机密标签缺乏保护效果。与无标签基线相比,添加“公开-允许共享”在描述上与更高的逐字信息流出相关,且具有强模型依赖性:Claude Sonnet 5表现强且一致(公开减无标签均值+0.800,覆盖全部10个场景;主要与Claude是否中继相关),某一GPT-5.6层级表现中等但受下限限制,另一层级表现小(中位数0),第三层级完全处于下限。这是单一配置中的关联,非因果或通用效应。代码、字节固定轨迹及离线分析流水线作为公开工件发布。
英文摘要
Safety properties assessed separately for Model Context Protocol (MCP) tool use and Agent2Agent (A2A) delegation need not describe behavior when one agent uses both. We measure one such behavior in a single controlled MCP-to-A2A configuration: a testbed drives a real-model host across a local MCP and a local A2A leg into an ordered event trace scored by exact deterministic rules (no LLM judge), one restricted decision per trial. In a pre-specified, frozen three-arm design, each of 10 record scenarios appears with a CONFIDENTIAL header, with no header, and with PUBLIC - OK TO SHARE; the six substantive record values are byte-identical across arms, and the outcome is verbatim occurrence of any of them in the outbound message. Four models x 3 arms x 4 repeats give 480 trials; the scenario is the unit of generalization, and we report the 10 scenario-level values (mean, median, sign counts), with no p-values or intervals. The confidential-minus-unlabeled contrast is inconclusive and floor-limited in every model (both arms at or near zero), so it does not show that confidential labels lack a protective effect. Adding PUBLIC - OK TO SHARE is descriptively associated with higher verbatim egress relative to the unlabeled baseline, with strong model dependence: strong and consistent for Claude Sonnet 5 (public-minus-unlabeled mean +0.800, all 10 scenarios; mostly an association with whether Claude relays at all), moderate but floor-limited for one GPT-5.6 tier, small (median 0) for another, and a complete floor for the third. This is an association in one configuration, not a causal or general effect. Code, byte-pinned traces, and the offline analysis pipeline are released as a public artifact.
Comments12 pages, 1 figure. Code and reproducibility artifacts: https://github.com/ArpanKumarM/agent-interop-bench/releases/tag/paper-v1