发表机构
University of Southern California (USC)(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过分析开源类型化决策模型,发现标签而非定义主导预测结果,定位问题源于提示渲染,并提出测试与缓解措施。
AI 中文摘要
类型化决策模型通过为调用者定义的多个选项中的每一个返回一个概率,来回答关于输入的固定问题。每个选项都带有一个简短的标签和一份书面定义,开发者在此定义中陈述模型应应用的规则。Jev为路由、审核和分诊引入了这一接口,随后出现了开源实现,并且每当语言模型通过给标签字符串打分而用作分类器时,都会发生同样的操作。我们研究了这些开源实现,其权重我们可以检查和修补,并询问概率是遵循定义还是标签。对标签的偏好我们称之为选项标签偏差。在四个开源权重类型化决策模型、从Qwen2.5骨干读取答案的三种方式、十一个分类任务以及PolicyBench(我们引入的一个合成路由套件,其中规则仅出现在定义中)中,答案大多来自标签。删除所有定义后,准确率保持不变(laya-td:0.8559对比0.8487),尽管这些定义本身支持0.7971的准确率,而将选项重命名为A和B会使准确率提高+0.1511 [+0.1377, +0.1646]。一个系统von不受影响,两个代码库在一个表达式上有所不同:laya将每个选项写成“{label}: {definition}”,而von只写定义。在未改变任何权重的情况下,双向更改该表达式,使所有三个laya检查点完全不变(+0.0000 [+0.0000, +0.0000]),并在von中产生了效果,当标签与其定义矛盾时,其准确率从0.8511降至0.2281。早期工作将此失败归因于这些模型使用的受限决策头而非文本解码器;我们的结果将其定位在提示渲染中。我们给出了一个两次调用测试,让从业者知道哪种情况适用于他们的模型,并衡量了四种缓解措施的价值。
英文摘要
A typed decision model answers a fixed question about an input by returning a probability for each of several caller-defined options. Each option carries a short label and a written definition, which is where a developer states the rule the model should apply. Jev introduced this interface for routing, moderation and triage, open implementations followed, and the same operation occurs whenever a language model is used as a classifier by scoring label strings. We study the open implementations, whose weights we can inspect and patch, and ask whether the probability follows the definitions or the labels. A preference for the label we call option-label bias. Across four open-weight typed decision models, three ways of reading an answer from a Qwen2.5 backbone, eleven classification tasks and PolicyBench, a synthetic routing suite we introduce in which the rule appears only in the definitions, the answer is mostly the labels. Deleting every definition leaves accuracy unchanged (laya-td: 0.8559 against 0.8487), although those definitions support 0.7971 on their own, and renaming the options to A and B raises accuracy by +0.1511 [+0.1377, +0.1646]. One system, von, is unaffected, and the two code bases differ in one expression: laya writes each option as "{label}: {definition}", while von writes only the definition. Changing that expression in both directions, with no weight changed, makes all three laya checkpoints exactly invariant (+0.0000 [+0.0000, +0.0000]) and creates the effect in von, whose accuracy falls from 0.8511 to 0.2281 when a label contradicts its definition. Earlier work attributed this failure to the constrained decision head these models use in place of a text decoder; our results locate it in the prompt rendering. We give a two-call test that tells a practitioner which case applies to their model, and measure what four mitigations are worth.