发表机构
University of the West of England(西英格兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对边缘设备上MCP风格工具调用的可靠性,提出CPU基准测试,评估五个Sub-2B模型,发现Qwen2.5-1.5B最优,并揭示资源与可靠性权衡。
AI 中文摘要
资源受限的单板计算机,包括Raspberry Pi、NVIDIA Jetson Nano、Arduino UNO Q、Orange Pi和LattePanda,推动了设备端小语言模型(SLM)智能体的发展,这些智能体减少了对云的依赖,改善了数据本地性,并能容忍间歇性连接。模型上下文协议(MCP)风格的工具调用要求不仅仅是流畅的生成:智能体必须输出机器可读的JSON,选择正确的工具,提供所有必需的参数,并避免非预期的操作。我们通过评估五个参数低于二十亿的开源模型——Phi-1.5、Pythia-1.4B、TinyLlama-1.1B-Chat、Qwen2.5-0.5B和Qwen2.5-1.5B——在100个涵盖天气查询、网络搜索、计算、电子邮件撰写和任务创建的提示上,采用贪心解码和核采样,建立了一个平台无关的CPU基线。一个恢复解析器去除Markdown围栏,提取花括号分隔的子字符串,并评分可解析性、工具名称正确性、参数完整性和值一致性。在此标准下,Qwen2.5-1.5B在贪心解码下达到75%,在采样下达到79%;Qwen2.5-0.5B在贪心解码下达到72%,但在采样下降至32%。Phi-1.5得分为0%;Pythia和TinyLlama最多达到7%。严格的事后审计发现,1000个原始响应中只有5个可直接解析为JSON,暴露出对输出恢复的近乎完全依赖。CPU资源探测显示,Qwen2.5-1.5B需要7,960 MiB内存和30.782秒的平均延迟;Qwen2.5-0.5B使用3,637 MiB和10.627秒,揭示了边缘部署中可靠性与资源之间的权衡。这些结果并未直接覆盖上述命名的板卡或完整的MCP实现。安全部署需要模式验证、约束生成、最小权限执行以及对重大行动的人工升级。
英文摘要
Resource constrained single-board computers including Raspberry Pi, NVIDIA Jetson Nano, Arduino UNO Q, Orange Pi, and LattePanda motivate on-device small language model (SLM) agents that reduce cloud dependence, improve data locality, and tolerate intermittent connectivity. Model Context Protocol (MCP)-style tool invocation demands more than fluent generation: an agent must emit machine-readable JSON, select the correct tool, supply all required arguments, and avoid unintended actions. We establish a platform-agnostic CPU baseline by evaluating five open-weight models below two billion parameters Phi-1.5, Pythia-1.4B, TinyLlama-1.1B-Chat, Qwen2.5-0.5B, and Qwen2.5-1.5B on 100 prompts spanning weather retrieval, web search, calculation, email composition, and task creation, under greedy decoding and nucleus sampling. A recovery parser strips Markdown fences, extracts brace-delimited substrings, and scores parseability, tool-name correctness, argument completeness, and value agreement. Under this criterion, Qwen2.5-1.5B achieves 75% (greedy) and 79% (sampling); Qwen2.5-0.5B achieves 72% (greedy) but drops to 32% under sampling. Phi-1.5 scores 0%; Pythia and TinyLlama reach at most 7%. A strict post-hoc audit finds only 5 of 1,000 raw responses directly parseable as JSON, exposing near-total dependence on output recovery. A CPU resource probe shows Qwen2.5-1.5B requires 7,960 MiB and 30.782 s mean latency; Qwen2.5-0.5B uses 3,637 MiB and 10.627 s, revealing a reliability-resource trade-off for edge deployment. These results do not cover the named boards directly or a full MCP implementation. Safe deployment requires schema validation, constrained generation, least-privilege execution, and human escalation for consequential actions.