AI 中文总结
针对 MCP 智能体框架连线后语义变化问题,提出差分测试方法,通过 18 个夹具在四个 Python 集成中发现 13 个分歧,强调回归测试需明确信息需求与读取方式。
AI 中文摘要
有效的模型上下文协议(MCP)消息并不能保证智能体框架能够保留下游软件所需的区分。我们提出了一种差分测试方法,该方法将固定的工具结果贯穿每个框架的公共接口,并检查显式的消费者需求。在四个固定的 Python 集成中应用于 18 个设计的测试夹具,该方法识别出 13 个独特的夹具-任务分歧,涉及结构化值、声明的错误和丰富内容。Google ADK 在其观察路径中满足每个主要契约;其他集成则显示出特定于接口的更改或执行失败。更改是否重要取决于消费者:将缺失的可选字段视为等同于 null 可以解释 OpenAI 严格丰富内容失败的大部分原因。在探索性重放中,解析 JSON 文本恢复了更多结构化值,但也返回了错误的值以及源中不存在的字段的值。文档化设置有助于处理冲突输出的压力案例,但无法恢复单独的结构化槽位。故障挑战暴露了预言机弱点,并在新案例上测试了其修复。该研究提供了来自受控案例的可复现、特定于接口的证据,而非生产失败率或模型行为测量。其实际意义在于,连线后回归测试需要同时指定所需信息和消费者如何读取该信息。
英文摘要
Valid Model Context Protocol (MCP) messages do not guarantee that an agent framework preserves the distinctions downstream software needs. We present a differential testing method that follows a fixed tool result through each framework's public interfaces and checks explicit consumer requirements. Applied to 18 designed fixtures in four pinned Python integrations, the method identifies 13 unique fixture-task divergences involving structured values, declared errors, and rich content. Google ADK satisfies every primary contract in its observed path; the other integrations show interface-specific changes or an execution failure. Whether a change matters depends on the consumer: treating absent optional fields as equivalent to null explains most of OpenAI's strict rich-content failures. In an exploratory replay, parsing JSON text recovers more structured values but also returns incorrect values and values for fields absent at the source. Documented settings help in conflicting-output stress cases without restoring a separate structured slot. Fault challenges expose an oracle weakness and test its repair on fresh cases. The study provides reproducible, interface-specific evidence from controlled cases, not production failure rates or model-behavior measurements. Its practical implication is that post-wire regression tests need to specify both the information required and how the consumer reads it.
Comments10 pages, 7 tables, 1 figure. Submitted to SE4AgenticAI 2026. Reproducibility artifact: https://doi.org/10.5281/zenodo.22784143