黑盒大推理模型不确定性量化的越狱方法
Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models
浏览论文内容
中文总结 AI 辅助
针对大推理模型黑盒不确定性量化中过度自信问题,提出提示级松弛算子与J4U越狱技术,显著提升校准效果,平均ECE降低达5倍。
中文摘要 AI 辅助
尽管大推理模型(LRMs)在复杂推理方面表现出色,但通过强化学习进行的对齐往往会导致系统性的过度自信。在生产环境中,由于可能无法获取logits,稳健的黑盒不确定性量化(UQ)对于可信度和安全性至关重要。针对LRMs的问答任务,我们表明现有的黑盒方法,如基于释义的自洽性和置信度言语化,相比简单的重复采样几乎没有提供改进,这表明对齐抑制了有用的输出变异性。我们引入了提示级松弛算子,通过近似使用更强KL正则化参数获得的最优策略的效果,来拓宽模型的有效输出分布,从而使其更接近参考模型。理论上,我们证明了松弛能够改善校准。我们提出了用于不确定性量化的越狱方法(J4U),这是一种源自越狱的技术,能够经验性地重现我们松弛理论所预测的行为特征。在3个数据集和4个LRM上,包括一个闭源生产模型,J4U相比重复采样的改进在多达6倍多的LRM-数据集-指标设置中达到了统计显著性,超过了我们所评估的最强黑盒UQ最先进基线,平均ECE降低幅度高达5倍。这些结果为黑盒LRM部署中的UQ提供了实用工具。
英文摘要
While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.
发表机构
- Petscraft
- Inria(法国国家信息与自动化研究所)
- Université Paris-Saclay(巴黎萨克雷大学)
- INSA CVL(法国国立应用科学学院中央-卢瓦尔河谷校区)
- LIFO(LIFO实验室)
机构由 AI 辅助整理,请以论文原文为准。