arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过智能体原生可复用工具原语实现大语言模型工具使用的工程化

Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives

Haibo Jin, Suijin Wang, Xucheng Yu, Haojing Luo, Haohan Wang

arXiv 2609.01736首次发表:更新:

发表机构

School of Information Sciences; University of Illinois at Urbana-Champaign; Starc Institute(信息科学学院; 伊利诺伊大学厄巴纳-香槟分校; 斯塔克研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出HEART框架,通过工具原语与ToolFace解决LLM工具使用的推理脆弱性与性能下降问题,在基准测试和现实任务中均实现性能提升并降低API成本。

AI 中文摘要

结合外部工具的大语言模型(LLM)在解决复杂现实任务中展现出卓越能力,但现有方法存在两大关键挑战:工具输出类型与API模式不兼容导致的脆弱多步及多轮推理,以及在工具目录规模较大时的性能下降。为解决这些问题,本文提出工具原语(Tool Primitives)设计,该设计以自然语言作为工具调用接口,替代僵化的基于API模式的调用方式,每个工具被封装为LLM接口,内部处理模式解析与执行,支持嵌套及多轮工具调用的自然工具间通信。基于工具原语,本文构建了包含25519个函数的集中式工具库ToolFace,LLM在推理时仅动态检索相关工具,无需在上下文环境中枚举原始API模式。为在复杂场景中可靠协调工具原语与ToolFace,本文进一步提出HEART框架,即基于智能体原生可复用工具原语的工具使用工程框架,包含规划器、路由器与验证器,三者协同支持动态工具调用规划、多步执行及反馈驱动的恢复。在五个基准测试上的实验表明,HEART平均比基于SFT的模型性能高10%,平均比GPT-5.4、Claude-4.6-Sonnet与Gemini-3.1-Pro高6%,且API成本最高降低85%;在50个现实任务上,HEART的任务完成率达84%,是三个前沿商业模型平均水平(22%)的3.8倍。

英文摘要

Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two key challenges: brittle multi-step and multi-turn reasoning caused by incompatible tool output types and API schemas, and performance degradation under large tool catalogues. To address these, we introduce \textbf{Tool Primitives}, a design that replaces rigid API schema-based invocation with natural language as the interface for tool calling, where each tool is wrapped with an LLM interface that handles schema resolution and execution internally, enabling natural inter-tool communication for nested and multi-turn tool calling. Building on Tool Primitives, we host \textbf{ToolFace}, a centralized repository of 25,519 functions from which LLMs dynamically retrieve only the relevant tools at inference time, eliminating the need to enumerate raw API schemas in context. To orchestrate Tool Primitives and ToolFace reliably in complex settings, we further propose \textbf{HEART}, a \textbf{H}arness \textbf{E}ngineering framework via \textbf{A}gent-native, \textbf{R}eusable \textbf{T}ool Primitives, comprising a Planner, Router, and Verifier that jointly support dynamic tool invocation planning, multi-step execution, and feedback-driven recovery. Experiments on five benchmarks demonstrate that HEART outperforms SFT-based models by $10\%$ on average and surpasses GPT-5.4, Claude-4.6-Sonnet, and Gemini-3.1-Pro by $6\%$ on average while reducing API cost by up to $85\%$. On 50 real-world tasks, HEART achieves $84\%$ task completion, $3.8\times$ the average of three frontier commercial models ($22\%$).

Comments21 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑