超越校准:类型化决策模型的概率是否遵循概率公理?
Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms?
浏览论文内容
中文总结 AI 辅助
本文提出无需标签的逻辑关联问题测试,发现类型化决策模型Jev和Qwen3.8-27B的概率违反概率公理,且不一致性随置信度变化,暴露了校准无法发现的系统偏差。
中文摘要 AI 辅助
类型化决策模型(如TypeSafe的Jev)以概率而非文本形式回答关于某个状态的已声明的是/否或多项选择问题,其评估报告准确性和校准性。这两者都不要求模型对逻辑相关问题给出的概率逐项匹配,而这正是基于这些概率行动的系统所需要的。我们通过一组无需标签的逻辑关联问题来测试这一属性,即一致性。对于来自ChaosNLI和PubMedQA的160个条目,每个条目有三个互斥的标签,我们询问标签是否为X、是否不是X、是否为其他两个标签之一,以及哪个标签适用。在480个否定对上,Jev对“标签是X”和“标签不是X”的概率平均偏离总和1达0.064(95%置信区间0.055至0.072)。使用官方BF16权重运行的Qwen3.8-27B,其首词概率偏离0.293,言语化概率偏离0.122。在两个系统给出相似概率的对上,与首词读数的差距持续存在,在没有双重否定标签的情况下,并且在平均Jev的重复调用后依然如此。Jev也不一致:其违规约为其重复噪声的五倍,并且它过度支持关于单个标签的陈述,因此其三个单标签概率平均总和为1.14。两个系统的失败方式也不同。Qwen3.8-27B的首词读数无论问题是否包含“不”,都低估标签的补集,在480对中拒绝一个陈述及其否定共196对,并且在其更自信的地方并未变得更一致,而Jev的违规集中在其答案不确定的地方。由于这些检查不需要标签,它们暴露了仅在比较问题形式时出现的偏差,以及条目内部的不一致性。
英文摘要
Typed-decision models such as TypeSafe's Jev answer a declared yes/no or multiple-choice question about a state with a probability instead of text, and their evaluations report accuracy and calibration. Neither requires that the probabilities a model gives to logically related questions fit together item by item, which is what a system that acts on those probabilities needs. We test this property, coherence, with a battery of logically linked questions that needs no labels. For 160 items from ChaosNLI and PubMedQA, each with three mutually exclusive labels, we ask whether the label is X, whether it is not X, whether it is one of the other two labels, and which label applies. On 480 negation pairs, Jev's probabilities for "the label is X" and "the label is not X" miss summing to one by 0.064 on average (95% CI 0.055 to 0.072). Qwen3.8-27B, run from its official BF16 weights, misses by 0.293 with first-token probabilities and by 0.122 with verbalized probabilities. The gap to the first-token readout persists on pairs where both systems give similar probabilities, without double-negation labels, and after averaging Jev's repeated calls. Jev is not coherent either: its violations are about five times its repeat noise, and it over-endorses statements about single labels, so that its three single-label probabilities sum to 1.14 on average. The two systems also fail differently. Qwen3.8-27B's first-token readout under-endorses the complement of a label whether or not the question contains "not", rejecting both a statement and its negation in 196 of 480 pairs, and it does not become more coherent where it is more confident, whereas Jev's violations concentrate where its answer is uncertain. Because the checks need no labels, they expose biases that appear only when question forms are compared, and inconsistencies within items.