发表机构
University of Pisa(比萨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对系统一模型概率校准缺失的问题,本文提出 Sys1Cal-v1 基准数据集,通过 Jev 的三种原语评估校准,发现 Jev 在二元决策中隐含第三真值“我不知道”,恢复后软准确率从 0.771 提升至 0.978。
AI 中文摘要
Jev 的出现标志着系统一模型时代的到来,这类基础模型返回带有概率分布的结构化决策,而非文本。除了低成本和高速度,Jev 的核心承诺是这些概率是经过校准的:这一说法并未得到任何公开测试的支持,现有的外部基准评估的是置信度校准,而非每个返回选项的概率是否具有正确的数值含义。为解决此问题,我们引入了 Sys1Cal-v1 数据集,其中包含关于命题 $A$ 的真/假问题,而 $A$ 的精确概率 $P(A)$ 通过构造已知。每个条目通过 Jev 的三种原语——Noul、Choice 和 Score——进行查询,并通过与真实分布的总变差距离进行评估,该距离可用于估计系统一模型的软准确率。我们通过评估 Jev 和 SemIf(一个开源的 Choice 风格基线)来展示 Sys1Cal-v1 作为基准数据集的实用性。然而,在这项工作中,我们更深入地聚焦于 Jev,研究其 Score 和 Choice 答案的校准。特别是,我们发现了一种奇特的行为,可以通过假设 Jev 抑制了第三种真值(超越真和假)来解释。换言之,在 Choice 答案中,$P(A)$ 和 $P(\ eg A)$ 被呈现为 $P(A)+P(\ eg A)=1$,而求和项中缺少 $P(U)\ eq0$。恢复 $P(U)$ 使得 Choice 答案的中位软准确率从 0.771 提升至 0.978,这表明即使在二元决策中,Jev 也倾向于用第三个选项回答:“我不知道”。
英文摘要
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition $A$ for which the exact probability $P(A)$ is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, $P(A)$ and $P(\neg A)$ are presented as if $P(A)+P(\neg A)=1$, while a term $P(U)\neq0$ is missing in the sum. Recovering $P(U)$ leads to an improvement of median soft accuracy in \texttt{Choice} answers from $0.771$ to $0.978$, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.