arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03966cs.AIcs.LG

接口诱导的轨迹审查

Interface-Induced Trajectory Censoring

Wenbo Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究发现工具调用率由模型-接口栈决定,通过实验揭示接口审查轨迹的效应,发布预检查工具捕获静默失败,为智能体评估提供关键改进。

中文摘要 AI 辅助

智能体评估报告的工具调用率来自服务栈,当模型生成格式正确的调用时,该数值可能为零,因为接口会在下游组件接收前审查轨迹。在BFCL v4自身数据上,固定权重、案例、解码方式和随机种子,仅改变服务适配器,同一模型的得分在0.00或0.96/0.19之间波动;针对聊天模板和解析器的2×2实验定位了该效应:两个主效应均为零,全部效应来自交互项——无组件存在缺陷,仅修复契约的某一方毫无作用。在tau-bench的115个交互式零售任务上,相同的适配器切换使服务器解析的调用数从0增至636,成功触发工具执行的任务数从0增至103。本探测实验在21倍规模的Qwen2.5-Coder上复现了该漏斗效应:各规模下服务器解析率均为0/100,而模型生成的格式正确调用率在32B规模时升至80/100(经人工裁决黄金标准校准后约为72)。在匹配的参数包下,跨可比规模范围,静默失败率维持在0-2,该预测在实验运行前已提交至代码库。Llama-3.1-8B将任务函数自身作为工具调用的比例为23%,设置strict:true标志后降至0。这种不匹配存在于训练循环内部,其后果具有规模依赖性:在verl的AgentLoop中,7B规模下115个生成结果中有45个包含完整调用,但均未被接受、执行或返回观察结果;1.5B规模下该零结果由多重因素导致,故报告两个规模的结果。评估阶段,修复适配器可恢复机制但未带来显著性能提升:解析率从0升至84,成功触发数从0升至9,通过率从53升至62(无统计学显著性)。本研究发布了一个98行的预检查工具,可捕获此处所有静默失败,观测到的工具调用率并非模型单独的属性,而是测量它的模型-接口栈的属性。

英文摘要

Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model is emitting well-formed calls: the interface censors the trajectory before anything downstream sees it. On BFCL v4's own data, executor and scorer, holding weights, cases, decoding and seeds fixed and changing only the serving adapter, the same model scores 0.00 or 0.96 / 0.19. A 2x2 over chat template and parser locates the effect exactly: both main effects are exactly zero and all of it sits in the interaction -- no component is defective, and repairing one side of the contract buys precisely nothing. On tau-bench's 115 interactive retail tasks the same swap moves server-parsed calls from 0 to 636 and tasks reaching any tool execution from 0 to 103. Our probe reproduces the funnel across a 21x scale range of Qwen2.5-Coder: the server parses 0/100 at every size while well-formed emitted calls rise to 80/100 at 32B (~72 after calibration against an adjudicated gold standard). Under a matched envelope, across a comparable scale span, the silent fraction stays at 0-2, a prediction committed to the repository before the run. Llama-3.1-8B's 23% rate of calling the task function itself as a tool falls to 0 under one strict:true flag. The mismatch reaches inside the training loop, and its consequence is scale-dependent: in verl's AgentLoop at 7B, 45 of 115 generations carry a complete call; 0 are accepted, 0 execute, 0 return an observation. At 1.5B the same zero is over-determined, so we report the two scales separately. At evaluation time, repairing the adapter restores the mechanism but not a significant outcome gain: parsing 0->84, rescues 0->9, pass rate 53->62 (n.s.). We release a 98-line preflight check that catches every silent failure here. The observed tool-call rate is not a property of the model alone; it is a property of the model-interface stack that measures it.

发表机构

  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑