发表机构
School of Mathematical and Computational Sciences, Massey University; RMIT University(梅西大学数学与计算科学学院; 皇家墨尔本理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究对比OpenClaw与NanoBot两款智能体系统,发现两者完整任务完成率无统计差异,NanoBot部分完成率更高但资源消耗更优,建议智能体AI评估需结合能力、资源及执行记录。
AI 中文摘要
向更自主的AI发展越来越依赖将语言模型与工具、记忆、状态管理及多步骤执行相结合的智能体系统,这些机制既影响任务能力也影响操作负担。我们将OpenClaw和NanoBot作为完整智能体系统进行对比,采用配对主基准测试及配对提示的更详细检测子集。在主基准测试中,OpenClaw的完整任务完成率为31%,NanoBot为25%,两者相差6个百分点,95%任务自举区间为-3至15个百分点,未统计出任一系统具备完整完成优势。在检测层,两个系统的完整完成率均为26%,而NanoBot在43%的提示上至少实现部分完成,OpenClaw仅为26%;83%的提示中OpenClaw耗时更长,且在所有提示上的记录峰值内存更高, wall time的几何均值比为2.98,峰值内存的几何均值比为19.44。在10个详细层提示中(至少一个系统实现部分或完整完成),NanoBot在8个提示上弱占优;在全部23个提示中,其18个占优案例里有10个为更经济的联合失败。两层证据的结果标签存在差异,表明智能体系统评估应将能力和资源测量与尝试级执行及评分来源关联。这些发现显示,向更自主AI发展的评估应通过经验证的任务完成情况、观测到的资源使用情况,以及将每个结果与其产生的执行关联的记录来开展。
英文摘要
Progress toward more autonomous AI increasingly depends on agentic systems that combine a language model with tools, memory, state management, and multi-step execution. These mechanisms shape both task capability and operational burden. We compare OpenClaw and NanoBot as complete agentic systems using a paired primary benchmark and a more detailed instrumented subset of paired prompts. In the primary benchmark, the rate of full task completion was 31% for OpenClaw and 25% for NanoBot, a six-percentage-point difference with a 95% task-bootstrap interval from -3 to 15 percentage points, providing no statistically established full-completion advantage for either system. In the instrumented layer, both systems achieved 26% full completion, while NanoBot reached at least partial completion on 43% of prompts compared with 26% for OpenClaw. OpenClaw took longer on 83% of prompts and had a higher recorded peak-memory value on every prompt, with geometric mean ratios of 2.98 for wall time and 19.44 for peak memory. Among the ten detailed-layer prompts on which at least one system achieved partial or full completion, NanoBot weakly dominated on eight; across all 23 prompts, however, ten of its eighteen dominance cases were cheaper joint failures. Outcome labels differ across the two evidence layers, showing why agent-system evaluation should connect capability and resource measurements to attempt-level execution and scoring provenance. These findings show that progress toward more autonomous AI should be evaluated through verified task completion, observed resource use and records linking each result to the execution that produced it.