arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30454cs.LG

审计系统-1模型在生物安全相关基准上的表现:校准、选择性预测与非生成模型中的排列不稳定性

Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Kimon Antonios Provatas, Ilias Georgakopoulos-Soares

AI总结:

该研究审计了一个商业系统-1模型在生物安全相关基准上的可靠性,发现其校准和选择性预测总体良好但在弱任务上退化,且对选项顺序敏感,而跨旋转平均概率可低成本提升准确性。

AI中文摘要:

非生成式“系统-1”模型在单次前向传播中返回结构化的概率决策,无需自回归解码,其推理成本仅为生成式模型的一小部分。这使得它们作为大型流水线中的低成本组件颇具吸引力,但其在生物安全相关任务上的可靠性尚未得到系统检验。我们审计了一个商业系统-1模型,在来自大规模杀伤性武器代理(WMDP)、一个释义鲁棒的WMDP-Bio变体以及六个LAB-Bench子任务的6020道多项选择题上,测量了准确性、校准、错误检测、选择性预测以及对答案选项呈现顺序的敏感性。准确性强烈依赖于任务。一旦正确解释供应商的不确定性字段,该模型的校准较为合理(合并期望校准误差为0.034),其最高概率能够区分正确与错误的预测(合并AUROC为0.820),尽管在较弱的任务上这两者均大幅下降。在答案选项的四种循环旋转下,37.4%的WMDP-Cyber项目获得了不同答案;一个使用字节相同重复调用的对照实验将大部分差异归因于选项顺序而非运行间变化。跨旋转平均概率使WMDP-Cyber准确性提高了3.8个百分点,且仅对低置信度项目应用该操作,便能以远低于平均所有项目的成本恢复大部分增益。

英文摘要:

Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larger pipelines, but their reliability on biosecurity-relevant tasks has not been systematically examined. We audit one commercial System-1 model on 6,020 multiple-choice items drawn from the Weapons of Mass Destruction Proxy (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, measuring accuracy, calibration, error detection, selective prediction, and sensitivity to the order in which answer options are presented. Accuracy is strongly task-dependent. Once the vendor's uncertainty field is correctly interpreted, the model is reasonably well calibrated (pooled expected calibration error 0.034) and its top-1 probability separates correct from incorrect predictions (pooled AUROC 0.820), though both degrade substantially on the weaker tasks. Under four cyclic rotations of the answer options, 37.4% of WMDP-Cyber items receive different answers; a control using byte-identical repeated calls attributes most of this to option order rather than run-to-run variation. Averaging probabilities across rotations improves WMDP-Cyber accuracy by 3.8 percentage points, and applying it only to low-confidence items recovers most of that gain at well under the cost of averaging every item.

补充信息

↑