发表机构
Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示本地服务栈(如Ollama、vLLM等)会显著影响编码智能体工具调用评估结果,提出将服务行为纳入评估协议并给出检查清单,以避免混淆模型真实能力。
AI 中文摘要
编码智能体必须发出有效的工具调用——即按照提供的模式对工具进行可解析的调用——之后测试平台才能执行其选择的动作。我们研究本地服务栈如何影响这一协议步骤,并表明测量结果可能取决于服务层,而非仅取决于模型行为。在Ollama中,默认的tools=请求按模型由静态模板标志门控:某些模型被接受并返回文本形式的调用,某些返回原生tool_calls,而Phi-3和Gemma-3在推理前被拒绝。在我们的测试平台中,拒绝和重试耗尽未被保留为结构化失败元数据,因此下游分析可能将其误分类为模型未调用,并天真地报告0%的保真度。在保留原生通道的同时添加文本工具列表,可恢复已接受模型的大部分测量保真度,而统一的文本协议会降低Llama-3.2的保真度,因为该模型具有原生工具调用支持。在Ollama、this http URL、vLLM和SGLang上的跨栈探测显示,对相同请求的处理方式不同。约束解码消除了解析失败,但可能导致不终止,且回合池化与逐实例估计的差异高达约55个百分点。我们最后提供一个清单,将服务行为视为评估协议的一部分。
英文摘要
A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavior alone. In Ollama, the default tools= request is gated per model by a static template flag: some models are accepted and return calls as text, some return native tool_calls, while Phi-3 and Gemma-3 are rejected before inference. In our harness, rejection and retry exhaustion are not preserved as structured failure metadata, so downstream analysis can misclassify them as model non-calls and naively report 0% fidelity. Adding a text tool list while retaining the native channel recovers much of the measured fidelity for accepted models, whereas a uniform text protocol reduces fidelity for Llama-3.2, which has native tool-call support. Cross-stack probes on Ollama, llama.cpp, vLLM, and SGLang show different handling of the same request. Constrained decoding removes parse failures but can induce non-termination, and turn-pooled versus per-instance estimates differ by up to about 55 points. We conclude with a checklist for treating serving behavior as part of the evaluation protocol.
Comments9 pages, 4 figures, 3 tables. Accepted at the 2nd Workshop for Research on Agent Language Models (REALM) @ EMNLP 2026