小型模型也能合理地进行自我置信度评估
Also Small Models Can Reasonably Self-Evaluate Their Confidence
- Technical University of Munich(慕尼黑工业大学)
- Siemens AG(西门子股份公司)
- LMU Munich(慕尼黑大学)
- Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究系统评估不同规模语言模型在问答任务中的自我评估置信度,发现模型规模和领域特异性对置信度可靠性影响不大,小型模型也能合理自我评估,适合资源受限部署。
AI中文摘要:
本研究系统评估了基于自我评估的不确定性量化方法,这些方法应用于不同规模的语言模型,并覆盖从通用到专业领域的问答任务。通过使用多种自我评估方法(模型对自身预测进行判断),我们考察了模型规模和领域特异性如何影响自我评估置信度信号的质量。结果显示,尽管随着模型变小和领域更专业,准确率可预期地下降,但自我评估置信度的可靠性在这两个维度上基本保持稳定。这种独立性意味着能力最强的模型未必是自我评估预测可靠性最佳的模型。这些发现表明,小型模型尽管准确率较低,仍能实现合理的自我评估置信度,从而使其适用于资源受限的部署场景。
英文摘要:
This study systematically evaluates self-evaluation-based uncertainty quantification across different language models of varying sizes on question-answering tasks spanning general to specialized knowledge domains. Using various self-evaluation methods where models judge their own predictions, we examine how model scale and domain specificity affect the quality of self-assessed confidence signals. Our results reveal that while accuracy predictably declines with smaller models and more specialized domains, the reliability of self-evaluated confidence remains largely stable across both dimensions. This independence means the most capable model is not necessarily the best at self-assessing prediction reliability. These findings suggest that smaller models can achieve reasonable self-assessed confidence despite lower accuracy, making them viable for resource-constrained deployments.