发表机构
National University of Singapore; Fudan University(新加坡国立大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过重命名选项名称实验,发现类型化决策模型受选项名称语义极性影响,导致决策排序逆转,但类型错误率仍为零。
AI 中文摘要
类型化决策模型是为模型输出直接被软件消费的场景而构建的。它们不是生成自由形式的文本,而是在预定义的选项集合上返回一个决策。通过构造,每个输出都符合所需的模式。然而,这种保证并不能告诉我们模型是否按预期解释选项。我们通过改变选项名称分配给评分标准的方式,研究了Jev和两个具有开放权重的类Jev模型。每个选项由一个选项名称和一个定义选项含义的文本评分标准组成。我们只改变分配给每个评分标准的选项名称;问题、状态、评分标准措辞和选项名称集合保持完全相同。在1200个具有任务特定评分标准的工作流决策中,将两个选项从0/1重命名为no/yes,每百个答案中改变了70.4个(95%置信区间:[67.6, 73.1]),并将AUC从0.94降至0.23,揭示了决策排序的系统性逆转,而非简单的不确定性。同样的操作对中性选项名称影响甚微。这种模式在所有4个谓词中一致,其影响至少比中性对照大7.4倍,并且随着选项数量的增加而增强。该影响还取决于读出几何结构:第二个模型家族对完整选项跨度进行平均池化,其翻转频率低4.1倍。托管模型表现出相同的行为:交换将AUC从0.8146降至0.5806,并产生比其测试-重测基线多24倍的答案翻转。相比之下,将选项名称替换为随机字符串会使所有模型家族回到中性对照状态,且不降低准确性。因此,失败取决于选项名称的语义极性,而非重命名操作本身。在所有条件下,类型错误率保持为0%,即使决策准确性大幅下降。
英文摘要
Typed decision models return structured results, but output-type correctness alone does not ensure that decisions follow explicit option definitions. Each option pairs a name with a definition that defines its intended meaning; the name, however, can provide a competing semantic cue. We study this conflict in Jev and two open-weight models by changing only the name-definition mapping, leaving the question, state, and the names and definition texts themselves unchanged. We measure decision flips at the level of the selected definition, rather than the returned name. On 1200 decision tasks with task-specific definitions, decision-flip rates are up to 70.4 pp higher with yes/no names than with the 0/1 control. This gap holds across all 4 binary decision rules. With yes/no names, reassignment also lowers their mean AUC from 93.8% to a below-chance 23.2%. In the binary evaluations, random strings used as option names yield mean flip rates close to those of neutral controls across all three models, with comparable balanced accuracy before reassignment. Together, these results support option-name polarity as a contributor to decision instability beyond reassignment alone. The type-error rate remains 0% throughout, showing that type-correct outputs can still fail to follow explicit option definitions.