修复,而非改进:分解工具调用弃权中的约束解码
Repair, Not Improvement: Decomposing Constrained Decoding in Tool-Call Abstention
浏览论文内容
中文总结 AI 辅助
该研究针对工具调用弃权问题,分解约束解码的作用,发现修复格式约束的效果优于改进,且两项预注册语言声明不成立。
中文摘要 AI 辅助
函数调用是近期约束生成研究明确排除的内容:该研究发现解码器对格式约束的贡献很小,随后在第7节警告不要将约束编码为正确性要求的情况进行推断,并将函数调用列为其中一种。工具弃权是该情况的极端体现:枚举(enum)仅保留答案的措辞并缩小答案集合,而弃权(不执行)调用任何工具是其首先会排除的情况。我们对这一被排除的情况进行了衡量。我们在一个字节完全相同的提示上设置三种条件,以区分语法的两项任务:它既确定生成的停止位置,也确定可生成的token。我们对0.6B到4B规模的开源权重模型在匹配的英语和韩语条目上进行评估,因此语言比较是在条目内部进行的。与无约束解码器相比,现有研究的对比结果在6个单元中有4个的弃权结果为负,且区间不包含0,最差为-29.5个百分点;没有任何单元的弃权结果为正且区间不包含0。总结果是符号相反的数值之和:最小模型在韩语上的停止标记造成-20.0的影响,枚举返回+19.5的影响,两者合计为-0.5。它修复的是形式:在698次弃权(不执行)中,545次没有可读答案,0次是评分者拒绝的判断。在需要工具的条目上,结果始终为正;弃权(不执行)是预注册的衡量指标,而汇总数值对干预更友好,因此转向弃权(不执行)会使结果变差而非变好。两项预注册的语言声明均不成立。
英文摘要
A tool-calling router has to pick the right tool when one applies and decline when none does. Restricting the decoder to a grammar over the tool names is the standard remedy for the first, and on small models it buys a large accuracy gain. Recent work separates the loss caused by asking for a format from the loss caused by enforcing it at decode time. The second is small, which made enforcement look nearly free. The same work declines to extend that to function calling, where a constraint decides which answers exist rather than how one is written. Declining to call anything is the answer it most easily removes, and the one a router can least afford to lose. A grammar decides which tokens may be emitted and where generation stops, and a two-condition design charges both to the restriction. We therefore run three conditions over one prompt: free generation, generation stopped at the first line, and both applied together. We evaluate open-weight models from 0.6B to 4B on the same items in English and Korean, comparing the languages item by item. The two-condition contrast is negative on abstention accuracy in four of six cells with intervals excluding zero, and positive in none, costing -29.5 points at worst. On the smallest model in Korean the stop costs -20.0 points, the restriction returns +19.5, and together they leave -0.5. What the restriction gives back is readable output, not judgment. Of the 698 abstentions it repairs, 545 had no readable answer at all and 0 were correct decisions the scoring rule rejected. On items that do need a tool the contrast is positive throughout, and abstention is reported first because it is the registered measure. Both preregistered claims about language fail: Korean does not lose more of the abstentions it holds without the constraint, and the removed mass does not explain what does.
发表机构
- Redrob
机构由 AI 辅助整理,请以论文原文为准。