arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34454cs.CL

当言语不足时:言语化推理与隐藏特征在LLM置信度估计中的迭代协同

When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yekun Xu, Ante Wang, Jingyi Ren, Xuanyi Chen, Weizhi Ma, Yang Liu

AI总结:

本文提出IPoET框架,通过迭代协同言语化推理与隐藏特征,提升LLM置信度估计,实证表明其优于现有基线。

AI中文摘要:

置信度估计对于开发可信的大型语言模型(LLM)至关重要,大多数方法遵循基于估计器或基于言语化的范式。尽管近期研究日益关注改进言语化的置信度自我报告,我们挑战了这种方法优于独立置信度估计器的普遍观点。我们的实证研究表明,专门的置信度估计器能显著优于言语化置信度,表明LLM的内部表示包含更丰富的置信度信号。基于这一发现,我们提出了迭代策略-估计器训练(IPoET),这是一个协同利用言语化推理轨迹和信息丰富表示的互补优势的框架。IPoET交替进行策略优化和估计器更新,将估计器得出的置信度反馈整合到策略学习中,并在新的策略rollout上刷新估计器。跨多个数据集以及Qwen和Llama骨干网络的实验表明,通过迭代利用更丰富的隐藏特征并适应不断演变的策略分布,IPoET在域内持续优于基于估计器和基于言语化的基线,并在所有域外指标上取得更优或相当的结果。更多细节,请参阅此https URL。

英文摘要:

Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpasses independent confidence estimators. Our empirical study shows that a dedicated confidence estimator can substantially outperform verbalized confidence, indicating that LLMs' internal representations contain richer confidence signals. Building on this finding, we propose Iterative Policy-Estimator Training (IPoET), a framework that synergizes the complementary strengths of verbalized reasoning traces and informative representations. IPoET alternates policy optimization with estimator updating, integrating estimator-derived confidence feedback into policy learning and refreshing the estimator on new policy rollouts. Experiments across diverse datasets and Qwen and Llama backbones demonstrate that, by iteratively exploiting richer hidden features and adapting to the evolving policy distribution, IPoET consistently outperforms both estimator- and verbalization-based baselines in-domain and achieves superior or comparable results across all out-of-domain metrics. For more details, refer to https://github.com/xyk829/ipoet.

↑