arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

校准是瓶颈:多轮工具调用的动作类诊断

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Kangjia Zhao, Jiajun Li, Haozhan Shen, Wei Chow, Linfeng Li, Hang Song, Lingdong Kong, Chen Zhi, Tiancheng Zhao, Songhua Liu, Jianwei Yin

arXiv 2609.00949首次发表:更新:

发表机构

Zhejiang University; Shanghai Jiao Tong University; National University of Singapore; Om AI Research; Binjiang Institute of Zhejiang University(浙江大学; 上海交通大学; 新加坡国立大学; Om AI研究院; 浙江大学滨江研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出面向动作类的多轮工具调用诊断框架,将失败分解为校准错误与执行失败,发现校准错误是被掩盖的关键瓶颈,且校准可通过上下文扰动异质性重塑,建议评估补充该诊断。

AI 中文摘要

多轮工具调用是大语言模型(LLM)智能体的核心评估场景。在公开工具调用基准上,开源权重模型的综合准确率已接近甚至超越闭源前沿模型。但该指标对多种不同的多轮场景取平均,掩盖了进展是否在这些场景间均衡分布。本文提出一种面向动作类的诊断框架,将多轮失败分解为两个正交模式:动作类校准错误与动作执行失败。该框架在四类动作空间(TOOL_CALL/调用工具、ASK/询问、REFUSE/拒绝、CONFIRM/确认)上运行,并引入自揭示上界Acc ≤ GAR(黄金动作召回率);两种模式表现为上界违反(Acc > GAR,暴露状态分级器对校准错误的掩盖)与上界大幅松弛(GAR >> Acc,将执行失败定位在TOOL_CALL类)。我们在多个多轮基准上的一组工具调用模型上验证了该框架。诊断结果显示,动作类校准错误是状态分级器无法察觉的重要失败模式。该差距夸大了大量工具训练模型族的表现,而我们的诊断可将其与具有上下文适配动作选择的模型族区分开。校准可仅通过上下文扰动重塑,但重塑效果存在异质性:同一扰动在不同模型族中使准确率向相反方向变动(同一场景下最高达+11.5个百分点与-21.0个百分点),且其效果进一步取决于扰动机制。本文认为,多轮工具调用评估应补充动作类诊断,以揭示模型在各场景下的实际行为。

英文摘要

Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However, this metric averages over many different multi-turn situations and obscures whether progress is balanced across them. We propose an action-class-oriented diagnostic framework that decomposes multi-turn failures into two orthogonal modes: action-class miscalibration and action-execution failure. The framework operates over a four-class action space (TOOL_CALL/ASK/REFUSE/CONFIRM) and introduces a self-revealing upper bound Acc <= GAR (Gold Action Recall); the two modes show up as bound violation (Acc > GAR, exposing state-grader masking of miscalibration) and large bound slack (GAR >> Acc, localizing execution failure within TOOL_CALL). We validate it on a panel of tool-calling models across multiple multi-turn benchmarks. Across our panel, the diagnostic reveals action-class miscalibration as a substantial failure mode the state grader cannot see. This gap inflates standing for heavily tool-trained families, which our diagnostic separates from families with context-appropriate action choice. Calibration is reshapable through context-only perturbations, but the reshape is heterogeneous: a single perturbation moves accuracy in opposite directions across families (up to +11.5 vs -21.0 pp on the same scenario), and its effect further depends on the perturbation mechanism. We argue that multi-turn tool-calling evaluations should supplement aggregate accuracy with action-class diagnostics that expose what the model actually does in each scenario.

CommentsAccepted to Findings of EMNLP 2026. Code: https://github.com/fbj2333/tool-calling-calibration

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑