参数优化:面向大语言模型工具调用的难度分级基准与探测引导训练
Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls
浏览论文内容
中文总结 AI 辅助
针对大语言模型工具调用参数填写关注度不足的问题,提出含PBT和PGR的探测引导框架,发布ParamBench基准,使5个开放模型在多基准上的参数生成平均精确匹配率从19.7%升至59.6%。
中文摘要 AI 辅助
大语言模型智能体的大部分能力来源于工具使用。现有工具使用研究主要聚焦于选择合适的工具和编排调用顺序,但正确填写工具调用的参数对于成功执行同样关键,却受到的关注少得多。在云网络等领域,即使是前沿模型,能正确完成的工具调用也不到一半。受近期分析显示大语言模型的隐藏状态包含关于模型预测的丰富信息的启发,我们发现,当模型生成参数值时,其隐藏状态包含强烈的正确性信号:一个简单的线性探测器能准确预测该值是否正确。基于这一观察,我们提出了一个统一的探测引导框架,包含两种互补方法:探测过滤的自举训练(PBT),该方法利用探测器过滤可靠的自生成调用以进行微调;探测引导重排序(PGR),该方法利用探测器在推理过程中选择更好的候选。为支持系统评估,我们发布了ParamBench,一个由真实云网络API构建的基准,它根据参数嵌套深度、跨参数依赖关系以及从先前调用推导值所需的推理,将每个实例分为五个难度级别。在ParamBench上对5个开放模型以及6个外部基准进行的大量实验表明,我们的方法大幅提升了参数生成性能,将平均精确匹配率从19.7%提高到59.6%。
英文摘要
Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool and orchestrating the order of calls. However, correctly filling the parameters of a tool call is equally critical for successful execution and has received far less attention. In domains such as cloud networking, even frontier models correctly complete fewer than half of tool calls. Inspired by recent analyses showing that LLM hidden states encode rich information about model predictions, we discover that while the model generates a parameter value, its hidden state contains a strong correctness signal: a simple linear probe can accurately predict whether the value will be correct. Based on this observation, we propose a unified probe-guided framework with two complementary approaches: probe-filtered bootstrapped training (PBT), which uses the probe to filter reliable self-generated calls for fine-tuning, and probe-guided reranking (PGR), which uses the probe to select better candidates during inference. To support systematic evaluation, we release ParamBench, a benchmark built from real cloud-network APIs that categorizes every instance into five difficulty levels according to parameter nesting depth, cross-parameter dependencies, and the reasoning required to derive values from earlier calls. Extensive experiments across 5 open models on ParamBench and 6 external benchmarks demonstrate that our method substantially improves parameter generation, raising the average exact match from 19.7% to 59.6%.