Pinocchio:黑盒语言模型的快速不确定性估计
Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models
浏览论文内容
中文总结 AI 辅助
Pinocchio是一种外部校准器,通过一次前向传播为黑盒API模型提供不确定性估计,无需访问内部状态,在七个LLM上训练后达到0.862 AUROC并零样本迁移至十三个模型。
中文摘要 AI 辅助
在大型语言模型(LLM)的高风险决策应用中,实践者不仅需要准确的LLM,还需要对其预测的不确定性估计。现有的LLM不确定性估计方法需要访问模型输出的对数概率或需要微调访问权限。然而,许多工业LLM产品使用闭源API模型,且许多此类API模型(如GPT)不返回对数概率,也可能不允许微调。我们引入了Pinocchio,一种外部校准器,用于估计黑盒API模型响应的正确性。它在七个LLM的响应上联合训练,在预测这些相同模型的留出响应正确性时达到0.862的AUROC,并展示了向八个组织的十三个未见模型的零样本迁移。我们的模型仅需一次前向传播即可生成不确定性估计,且无需访问目标模型的logits、权重或内部状态。一个轻量级仅文本0.8B检查点与我们的最大模型的AUROC相匹配。我们发布了代码,只需额外两行代码即可为现有仓库添加不确定性估计。
英文摘要
In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions. Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. However, many industrial LLM products use closed-source API models, and many such API models like GPT do not return log-probabilities and may not allow fine-tuning. We introduce Pinocchio, an external calibrator that estimates the correctness of responses from black-box API models. Trained jointly on responses from seven LLMs, it achieves 0.862 AUROC predicting the correctness of held-out responses from those same models, and shows zero-shot transfer to thirteen unseen models across eight organizations. Our model needs only a single forward pass to generate an uncertainty estimate and requires no access to the target model's logits, weights, or internal states. A lightweight text only 0.8B checkpoint matches our largest model's AUROC. We release code for adding uncertainty estimation to existing repos in only two additional lines of code.
发表机构
- Ritual AI
- Fudan University(复旦大学)
- Columbia University(哥伦比亚大学)
- University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。