发表机构
University of Montana(蒙大拿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对小型语言模型,探究口头表述置信度支撑风险可控 deferral 的可行性,通过理论推导与实验验证校准的局限性及认证式 deferral 的效果,修复了 TruthfulQA 的答案排序伪影。
AI 中文摘要
小型开放权重语言模型越来越多地运行在私密、离线且对成本敏感的场景中,其部署的关键问题不仅是模型会给出什么答案,还包括何时应将决策 defer 至人类。我们研究口头表述的置信度是否能支持风险可控的 deferral,对来自三个系列的 11 个指令微调模型(参数规模为 0.5B 至 14B)在 ARC-Challenge 和 TruthfulQA 数据集上进行评估,共完成 25168 次本地预测。三项理论结果界定了校准所能提供的内容:严格单调校准可保留风险-覆盖前沿和错误检测 AUROC;温度缩放无法校准那些置信度始终高于 0.5 但准确率却低于 0.5 的模型;Clopper-Pearson 过程在独立同分布部署假设下,可将包含 200 个问题的校准集转换为有限样本风险证书。实验方面,22 个模型-任务对中有 8 个达到了温度缩放不可行性下限,与预测边界的偏差在 1 个百分点以内;Platt 缩放可将 ECE 降低至低至 0.02,但在 20%风险预算下仅 3 个模型-任务对获得认证式自主权限,在 10%风险预算下则无任何模型-任务对获得该权限。我们还识别并修复了 TruthfulQA 多项选择形式中的答案排序伪影。校准赋予置信度语义;认证式 deferral 决定了小型模型何时可安全使用。
英文摘要
Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human. We study whether verbalized confidence can support risk-controlled deferral, evaluating eleven instruction-tuned models from three families, 0.5B to 14B parameters, on ARC-Challenge and TruthfulQA with 25,168 local predictions. Three theoretical results delimit what calibration can provide: strictly monotone calibration preserves the risk-coverage frontier and error-detection AUROC; temperature scaling cannot calibrate models whose confidence stays above one half while accuracy falls below it; and a Clopper-Pearson procedure converts a 200-question calibration set into a finite-sample risk certificate under an i.i.d. deployment assumption. Empirically, eight of 22 model-task pairs hit the temperature-scaling infeasibility floor within one percentage point of the predicted bound. Platt scaling reduces ECE to as low as 0.02, yet certified autonomy at a 20% risk budget is granted to only three model-task pairs and to none at 10%. We also identify and repair an answer-ordering artifact in the multiple-choice form of TruthfulQA. Calibration gives confidence semantics; certified deferral determines when small models are safe to use.
CommentsAccepted at MIWAI 2026 (The 19th International Conference on Multi-disciplinary Trends in Artificial Intelligence), to appear in Springer LNAI