arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从推理中估计不确定性:对大语言模型中多语言和跨语言MCQA性能的大规模研究

Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico

arXiv 2607.06327首次发表:更新:

发表机构

INSAIT; Amazon(萨格勒布信息与智能技术研究所; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究对22种语言的不确定性估计方法进行大规模评估,使用两个人工整理的问答数据集,比较不同模型和架构下的UE方法。发现促使模型用英语推理可提升UE性能,缩小不同资源语言间的性能差距,且UE方法选择取决于模型规模,还给出了选择性预测阈值选择的分析。

AI 中文摘要

不确定性估计(UE)能使基于大语言模型的系统识别何时弃权,但现有研究主要集中在英语上。我们首次对22种语言的UE方法进行大规模评估,涵盖高、中、低资源环境。使用两个人工整理的问答数据集,我们比较了不同模型大小和架构下的开箱即用和闭箱UE方法(共9种),同时引出长格式推理,避免可能引入评估噪声的大语言模型作为评判和基于嵌入的评分。我们报告了三个主要的可操作发现。首先,促使模型用英语推理,同时保持低资源语言的问题,能显著提高UE性能,这表明对低资源语言的理解基本完好,可靠性瓶颈在于生成而非理解。其次,促使模型用英语推理缩小了低资源和高资源语言之间的UE性能差距,表明生成语言比问题语言更重要。第三,UE方法的选择应取决于模型规模:在较小规模下,开箱即用的基于概率的方法优于其他方法;在较大规模下,闭箱自我表述的不确定性更优。最后,我们提供了选择性预测阈值选择的分析,为多语言环境下校准弃权提供指导。

英文摘要

Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-as-a-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we find that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting that comprehension of low-resource languages is largely intact, and that the reliability bottleneck lies in generation rather than understanding. Second, prompting models to reason in English closes the UE performance gap between low and high-resource languages, demonstrating that generation language matters more than the question language. Third, the choice of UE method should depend on model scale: at smaller scales, open-box probability-based methods outperform alternatives; at larger scales, closed-box self-verbalized uncertainty becomes superior. Finally, we provide an analysis of threshold selection for selective prediction, offering guidance on calibrating abstention in multilingual settings.

CommentsAccepted at Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑