How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains
大型推理模型的置信度估计有多可靠?对高风险领域的系统基准测试
AI总结 本文通过系统基准测试,评估了大型推理模型置信度估计方法的可靠性,发现基于文本的编码器在歧视方面表现最佳,而结构感知模型在校准方面表现最佳,揭示了当前方法的局限性。
Comments Accepted to the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026) main conference