ProbPlug:用于LLM二分类可靠置信度的插件不确定性网络
ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification
浏览论文内容
中文总结 AI 辅助
本文提出ProbPlug,一种轻量级插件网络,利用冻结LLM的内部token特征和自注意力机制,在不修改基础模型的情况下提供可靠置信度估计,提升二分类性能并具有强泛化能力。
中文摘要 AI 辅助
大型语言模型(LLMs)在广泛的分类场景中取得了强劲的性能,但其预测的可靠性仍然是部署于高风险场景中的主要障碍。尽管针对LLMs的置信度估计已被广泛研究,但基于LLM的分类任务的置信度校准仍未得到充分探索。我们提出了ProbPlug,一种用于基于LLM的二分类的轻量级置信度估计框架,它利用从冻结的LLM中提取的内部token特征来预测输出是否正确。ProbPlug采用自注意力模块来聚合隐藏表示,并且可以在不修改基础模型的情况下集成到原始推理流程中。在涉及基于文本和多模态大型模型的多个任务上的实验表明,ProbPlug提供了更可靠的置信度估计,以可忽略的额外开销提升了分类性能,并展现出跨任务的强泛化能力。这些结果表明,ProbPlug可作为基于LLM的分类中置信度估计的实用解决方案。我们的代码已在Github上公开。
英文摘要
Large language models (LLMs) have achieved strong performance across a broad range of classification settings, yet the reliability of their predictions remains a major obstacle to deployment in high-stakes scenarios. Although confidence estimation for LLMs has been widely studied, confidence calibration for LLM-based classification remains underexplored. We introduce ProbPlug, a lightweight confidence estimation framework for LLM-based binary classification, which predicts whether an output is correct using internal token features extracted from a frozen LLM. ProbPlug employs a self-attention module to aggregate hidden representations and can be integrated into the original inference pipeline without modifying the base model. Experiments across multiple tasks involving both text-based and multimodal large models show that ProbPlug provides more reliable confidence estimates, improves classification performance with negligible additional overhead, and exhibits strong generalization across tasks. These results indicate that ProbPlug serves as a practical solution for confidence estimation in LLM-based classification. Our code is publicly available at Github.
发表机构
- Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司)
- Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
- Hubei University of Technology(湖北工业大学)
机构由 AI 辅助整理,请以论文原文为准。