arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11392cs.AIstat.AP

TypedBench:面向第一类决策模型的校准、措辞敏感性及成本基准

TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models

Rahul Sharma, Andrew B. Ducan, Gaétan Marceau Caron, Sebastian J. Vollmer

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员构建TypedBench基准评估第一类决策模型,发现托管模型受措辞影响且置信不足,解码器路由准确但随选项增多变慢,需联合评估模型多方面性能。

中文摘要 AI 辅助

第一类模型通过非生成式接口输出针对类型化答案(如分类选择、有序等级或二元结果)的校准概率,软件可通过阈值、成本加权选择及升级规则基于这些概率采取行动。因此,若这些概率校准不当或受措辞影响,软件可能采取意外行动,使人类操作员无文本依据可检查。当前评估大多报告公开分类数据集上的准确率和校准度,但缺乏明确参考。我们提出TypedBench,这是一个基于7个策略标记生成器和9个评估套件构建的、针对Jev等类型化决策模型的基准。我们报告了不同释义下的准确率中位数和范围,以及相对于匹配的完美校准预测器的有限样本噪声基底的校准误差。我们通过选择性预测、有序恰当评分规则及非对称成本矩阵下的实际成本评估概率质量。我们在相同项目上评估了一个托管模型、一个开放编码器及一系列参数规模为0.8B至9B的开放解码器。托管模型遵循所述策略,但受措辞影响且系统性地置信不足;在非对称成本下,使用其概率可能比采用其最高置信度答案更差。解码器路由准确,但随着选项或问题增加而变慢;解码器在策略问题上准确率最低,且随选项增多而下降。总体而言,类型化决策模型必须联合评估策略 adherence、措辞鲁棒性、概率质量及诱导的决策结果。

英文摘要

System One models output calibrated probabilities over typed answers such as categorical choices, ordinal levels, or binary outcomes, via a non-generative interface. Software can act on these probabilities through thresholds, cost-weighted choices, and escalation rules. Consequently, if these probabilities are miscalibrated or wording-sensitive, the software ma take unintended actions leaving human operators with no textual rationale to inspect. Current evaluations largely report accuracy and calibration on public classification datasets without a clear reference. We present TypedBench, a benchmark for typed decision models like Jev built from seven policy-labelled generators and nine evaluation suites. We report accuracy as median and range across paraphrases, and calibration error relative to the finite-sample noise floor of a matched, perfectly calibrated predictor. We assess probability quality through selective prediction, ordinal proper scoring rules, and realised cost under asymmetric cost matrices. We evaluate a hosted model, an open encoder, and a family of open decoders spanning 0.8B-9B parameters on identical items. The hosted model follows the stated policy but is wording-sensitive and systematically underconfident; under asymmetric costs, using its probabilities can be worse than taking its top answer. Decoders route exactly and are slow as options or questions are added. The decoder is least accurate on policy questions and degrades with more options. Overall, typed decision models must be evaluated jointly on policy adherence, wording robustness, probability quality, and induced decision outcomes.

发表机构

  • DFKI GmbH(德国人工智能研究中心)
  • RPTU Kaiserslautern(凯泽斯劳滕工业大学)
  • Imperial College London(伦敦帝国学院)
  • Mila - Quebec AI Institute(米拉-魁北克人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑