发表机构
Hazel Mak(Hazel Mak)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过对比五种工具接口,发现仅用 bash 在企业任务上优于专用工具,得分提升显著且 token 消耗更少,为不同安全场景下的工具选择提供了实证依据。
AI 中文摘要
在本研究中,我们考察了通用 shell 是否能在企业任务上胜过专用工具。基于 shell 的代理在编码方面已展现出强劲表现,但企业工作还涉及在应用程序和服务之间切换、与同事协作以及进行专业分析。我们使用 Opus-4.8 和 GPT-5.5 在 TheAgentCompany 和 APEX-Agents 上比较了五种工具接口:类型化工具、类型化工具加 bash、仅 bash、带代理合成持久工具的 bash,以及程序化工具调用(PTC),后者运行其操作被限制在类型化工具目录中的程序。仅 bash 在两个基准上都优于类型化工具,在 TheAgentCompany 上得分提高 21.8-24.5 个百分点,在 APEX-Agents 上提高 4.8-7.4 个百分点,同时使用的总 token 减少 19-72%。在 bash 上添加类型化工具或持久工具合成并未产生可检测的合并得分增益。PTC 使用的 token 比直接类型化调用少,任务性能大致相似,但在质量和成本效率上通常不如仅 bash。对于企业从业者而言,这些结果支持在任意执行可被隔离时采用仅 bash,而在安全或合规策略要求固定工具目录时采用 PTC。
英文摘要
In this study, we examine whether a general shell can outperform specialized tools on enterprise tasks. Shell-based agents have shown strong results in coding, but enterprise work also involves moving between applications and services, coordinating with coworkers, and performing professional analysis. We compare five tool interfaces on TheAgentCompany and APEX-Agents using Opus-4.8 and GPT-5.5: typed tools, typed tools plus bash, bash alone, bash with persistent agent-synthesized tools, and programmatic tool calling (PTC), which runs programs whose actions are restricted to a typed tool catalog. Bash alone outperforms typed tools on both benchmarks, improving score by 21.8-24.5 pp on TheAgentCompany and 4.8-7.4 pp on APEX-Agents while using 19-72% fewer total tokens. Adding typed tools or persistent tool synthesis to bash produces no detectable pooled score gain. PTC uses fewer tokens than direct typed calls with broadly similar task performance, but generally underperforms bash alone in both quality and cost efficiency. For enterprise practitioners, these results favor bash alone when arbitrary execution can be isolated and PTC when security or compliance policies require a fixed tool catalog.
Comments13 pages, 7 figures, 12 tables