arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09021cs.AI

并非每次调用都需要前沿模型:已部署智能体家居自动化系统中按调用点的小语言模型评估

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

  • University of West Attica(西阿提卡大学)
  • Waldiez PC
  • ThinGenious PC

机构由 AI 辅助整理,请以论文原文为准。

Panagiotis Kasnesis, Christos Chatzigeorgiou, Lazaros Toumanidis, Amalia Contiero Syropoulou

AI总结:

本研究评估9个模型在智能体家居自动化系统五个调用点的性能,发现能力排序因调用点而异,路由最佳本地模型可达91.8%准确性,且仅托管生成性调用点可节省72%成本。

AI中文摘要:

智能体系统会发出几种结构上不同类型的LLM调用。它路由意图、分类动作、将语言基于设备注册表、规划多智能体流水线,并编写这些流水线运行的Python代码。这些调用点的难度相差一个数量级,然而在实践中,一个为最难调用点选择的单一模型服务于所有调用点。在这项工作中,我们评估了从0.8B到前沿托管模型的9个模型,在已部署的开源家居自动化框架(Wactorz)的五个调用点上,使用其未修改的生产提示和两个真实的Home Assistant安装(280个案例,2520次评分调用)。我们发现,能力在每个调用点的排序并不相同,且更大的模型并非普遍更好:一个4B模型在其2B兄弟模型在接地执行上表现更差。配对测试显示,在五个调用点中的四个,最佳本地模型与两个托管模型在统计上无显著差异。只有代码生成将它们区分开来,对抗一个小型托管模型(p = 0.039)以及一个前沿模型(p = 0.002)。聚合准确性也掩盖了特定于执行的安全失败,其中小模型以退化方式解决准确性/拒绝权衡:一个模型(Gemma4 E2B)在站点不拥有的设备请求上执行了87.2%的操作,而另一个模型拒绝其收到的每个请求。将每个调用点路由到其最佳本地模型达到91.8%,而95.4%无需每次调用成本。在由用户评判的实时部署中,仅托管两个生成性调用点与托管所有调用点匹配(39/43对39/43),花费仅为28%,且基准预测的执行差距恰好出现在二十六个案例中的一个。基准、测试框架和所有记录在此https URL发布。

英文摘要:

An agentic system issues several structurally different kinds of LLM calls. It routes intent, classifies actions, grounds language in a device registry, plans multi-agent pipelines and writes the Python code those pipelines run. The difficulty of these call sites varies by an order of magnitude, yet in practice a single model, chosen for the hardest site, serves all of them. In this work, we evaluate 9 models from 0.8B to a frontier hosted model across the five call sites of a deployed open-source home-automation framework (Wactorz), using its unmodified production prompts and two real Home Assistant installations (280 cases, 2520 scored calls). We find that capability is not ordered the same way at every site, and that larger models are not uniformly better: one 4B model is worse than its 2B sibling at grounded actuation. Paired testing shows the best local model to be statistically indistinguishable from both hosted models at four of five sites. Only code generation separates them, against a small hosted model (p = 0.039) as well as a frontier one (p = 0.002). Aggregate accuracy also hides a safety failure specific to actuation, where small models resolve the accuracy/refusal trade-off in degenerate ways: one model (Gemma4 E2B) actuates on 87.2% of requests for devices the site does not own, while another refuses every request it receives. Routing each site to its best local model reaches 91.8% against 95.4% at no per-call cost. In a live deployment judged by a user, hosting only the two generative sites matches hosting everything (39/43 against 39/43) for 28% of the spend, and the actuation gap the benchmark predicted appears as exactly one case in twenty-six. Benchmark, harness and all records are released at https://github.com/waldiez/slm-callsite-eval.

补充信息

↑