Guard模型在基础模型不确定的输入上过度自信
Guard Models Are Overconfident Where Base Models Are Uncertain
浏览论文内容
中文总结 AI 辅助
本文发现Guard模型在基础模型不确定的输入上过度自信,对抗攻击使其校准度大幅下降,但基础模型的不确定性信号仍可用来识别错误,揭示了Guard置信度与基础模型不确定性之间的不匹配。
中文摘要 AI 辅助
Guard模型用作安全分类器,其置信度分数驱动下游审核决策。我们评估了五个Guard模型在提示分类上的表现,发现尽管几个模型在干净输入上几乎校准,但对抗性攻击使它们的校准度降低了一个数量级,将假阴性转变为高置信度错误,与正确检测难以区分。将每个Guard与其对应的基础语言模型(LM)进行比较,我们发现不确定性信号通常仍然可用,基础模型通常在Guard失败的相同输入上表达不确定性。分层分析将这种Guard与基础模型的差异定位到较后的层,其中Guard模型表现出更尖锐的安全/不安全分离和较低秩表示,而对抗性有害输入更接近干净安全区域。这些发现凸显了攻击下Guard置信度与基础模型不确定性之间的不匹配。
英文摘要
Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions. We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections. Comparing each guard with its corresponding base LM, we find that uncertainty signals often remain available, with the base model typically expressing uncertainty on the same inputs where the guard fails. Layer-wise analyses localize this guard-base divergence to later layers, where guard models exhibit sharper safe/unsafe separation and lower-rank representations, while adversarial harmful inputs lie closer to the clean-safe region. These findings highlight a mismatch between guard confidence and base model uncertainty under attack.
发表机构
- DATUMO INC.(DATUMO公司)
机构由 AI 辅助整理,请以论文原文为准。