arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35350cs.AIcs.CLcs.LG

黑盒大推理模型不确定性量化的越狱方法

Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models

Lucas Biechy, Cédric Eichler, Adrien Boiret, Nicolas Anciaux

首次发表
浏览论文内容

中文总结 AI 辅助

针对大推理模型黑盒不确定性量化中过度自信问题,提出提示级松弛算子与J4U越狱技术,显著提升校准效果,平均ECE降低达5倍。

中文摘要 AI 辅助

尽管大推理模型(LRMs)在复杂推理方面表现出色,但通过强化学习进行的对齐往往会导致系统性的过度自信。在生产环境中,由于可能无法获取logits,稳健的黑盒不确定性量化(UQ)对于可信度和安全性至关重要。针对LRMs的问答任务,我们表明现有的黑盒方法,如基于释义的自洽性和置信度言语化,相比简单的重复采样几乎没有提供改进,这表明对齐抑制了有用的输出变异性。我们引入了提示级松弛算子,通过近似使用更强KL正则化参数获得的最优策略的效果,来拓宽模型的有效输出分布,从而使其更接近参考模型。理论上,我们证明了松弛能够改善校准。我们提出了用于不确定性量化的越狱方法(J4U),这是一种源自越狱的技术,能够经验性地重现我们松弛理论所预测的行为特征。在3个数据集和4个LRM上,包括一个闭源生产模型,J4U相比重复采样的改进在多达6倍多的LRM-数据集-指标设置中达到了统计显著性,超过了我们所评估的最强黑盒UQ最先进基线,平均ECE降低幅度高达5倍。这些结果为黑盒LRM部署中的UQ提供了实用工具。

英文摘要

While Large Reasoning Models (LRMs) excel at complex reasoning, alignment through reinforcement learning often induces systemic overconfidence. In production environments, where logits may be unavailable, robust black-box uncertainty quantification (UQ) is essential for trustworthiness and safety. Focusing on question-answering for LRMs, we show that existing black-box methods, such as paraphrase-based self-consistency and confidence verbalization, offer little to no improvement over simple repeated sampling, suggesting that alignment suppresses useful output variability. We introduce prompt-level relaxation operators that broaden the model's effective output distribution by approximating the effect of an optimal policy obtained with a stronger KL-regularization parameter, hence closer to the reference model. Theoretically, we demonstrate that relaxation improves calibration. We propose Jailbreak for Uncertainty (J4U), a jailbreak-derived technique for UQ that empirically reproduces the behavioral signatures predicted by our relaxation theory. Across 3 datasets and 4 LRMs, including a closed-source production model, J4U's improvement over repeated sampling achieves statistical significance in up to 6 times more LRM-dataset-metric settings than the strongest black-box UQ state-of-the-art baseline we evaluate, with average ECE reductions up to 5 times larger. These results provide a practical tool for UQ in black-box LRM deployment.

发表机构

  • Petscraft
  • Inria(法国国家信息与自动化研究所)
  • Université Paris-Saclay(巴黎萨克雷大学)
  • INSA CVL(法国国立应用科学学院中央-卢瓦尔河谷校区)
  • LIFO(LIFO实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑