arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一致但校准错误:评估大语言模型在自然语言风险沟通中的局限性

Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language

Diego Cerda-Mardini, Sarath Chandar, Sreenath Madathil

arXiv 2607.03882首次发表:更新:

发表机构

Faculty of Dental Medicine and Oral Health Sciences, McGill University; Chandar Research Lab, Polytechnique Montréal; Mila – Québec Artificial Intelligence Institute; Département de Génie Informatique et Génie Logiciel (GIGL), Polytechnique Montréal(麦吉尔大学牙医学与口腔健康科学学院; 蒙特利尔理工大学钱达尔研究实验室; 魁北克人工智能研究所米拉; 蒙特利尔理工大学计算机科学与软件工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型在自然语言中传达概率信息的可靠性,在两阶段预测管道中评估九个模型,发现其一致但校准错误,提供预计算统计量未解决问题,表明当前模型非可靠零样本风险沟通工具。

AI 中文摘要

大语言模型越来越多地被用作人工智能生成输出的事后解释器,但它们是否能在自然语言中可靠地传达概率信息仍不清楚。为了使这个角色可行,模型必须为相同的输入生成相同的语言描述,并选择准确反映潜在数值量大小的描述。我们在一个两阶段预测管道中评估九个大语言模型是否满足这些要求,其中上游模型产生了以可能性和不确定性为特征的概率输出,大语言模型的任务是为每个输出选择合适的语言描述符。我们通过从由其模式和先验样本大小参数化的贝塔分布中采样来模拟上游模型的预测。然后,我们在六种领域背景和十种温度设置下提示大语言模型解释这些预测,并重复每个实验十次。我们发现大语言模型通常是一致的,但校准错误,在不确定性任务上的表现明显弱于可能性任务。为模型提供预先计算的汇总统计量(模式和先验样本大小)降低了对上下文框架的敏感性,但没有解决潜在的校准错误,这表明瓶颈在于语言表达步骤本身。这些发现表明,当前的大语言模型还不能构成用于概率预测的可靠零样本独立风险沟通工具。

英文摘要

LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language. For this role to be viable, models must produce identical verbal descriptions for identical inputs, and select descriptions that accurately reflect the magnitude of the underlying numerical quantities. We evaluate whether nine LLMs meet these requirements within a two-stage prediction pipeline, in which an upstream model has produced probabilistic outputs characterized by their likelihood and uncertainty, and LLMs are tasked with selecting an appropriate verbal descriptor for each. We simulate predictions from an upstream model by taking samples from a Beta distribution parameterized by its mode and prior sample size. We then prompt LLMs to explain these predictions under six domain contexts and with ten temperature settings, and repeating each experiment ten times. We find that LLMs are generally consistent but miscalibrated, with substantially weaker performance on uncertainty than on likelihood tasks. Providing models with precomputed summary statistics (mode and prior sample size) reduced sensitivity to contextual framing but did not resolve the underlying miscalibration, suggesting that the bottleneck resides in the verbalization step itself. These findings indicate that current LLMs do not yet constitute reliable zero-shot standalone risk communication tools for probabilistic predictions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑