arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从可观察的语言模型概率校准语义不确定性

Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

Matthew F. Dixon

arXiv 2607.17447首次发表:更新:

发表机构

Artificial Intelligence Finance Institute (AIFI); Quiota LLC(人工智能金融研究所; Quiota公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何从语言模型概率校准语义不确定性,引入语义映射方法,经测试该方法在处理专业市场文本和模拟时,语言衍生概率表现良好,能恢复后验且稳定,语义映射将问题转化为可测试统计问题并产生可审计后验估计。

AI 中文摘要

语言模型生成单词概率,但专业决策需要对诊断、假设或操作条件等有意义状态的不确定性。模型打印的数值置信度无法确保可靠性。我们引入语义映射,它是从语言响应概率到声明状态概率的预设、可测试桥梁,通过半参数推断隐状态。参考模型定义目标后验,语言模型提供响应条件分布,留出法校准两者。我们推导后验误差界限及相关条件。通过对专业市场文本和精确后验的控制模拟测试,语言衍生概率表现优于打印数值概率,能恢复后验且不确定性覆盖有效,释义下稳定,对证据变化响应恰当。语义映射将一般问题转化为可测试统计问题,在接受条件满足时产生可审计后验估计,为审计分类等提供模板。

英文摘要

As generative artificial intelligence enters scientific and professional work, its uncertainty must be defined on the states that matter for inference and decision-making. Language models assign probabilities to words, whereas applications require uncertainty over meaningful states such as diagnoses, hypotheses or operational conditions. We introduce a \emph{semantic map}: a prespecified, testable bridge from probabilities over verbal responses to a posterior over declared finite states. The language distribution remains unrestricted; held-out calibration connects it to a reference posterior. We derive posterior-error bounds and conditions for existence, conditional uniqueness, presentation stability and stable inverse recovery. This distinction matters because language probabilities depend on prompt wording, while the target posterior should not change under information-equivalent rewording. Experiments use professional market text compiled from Federal Reserve economic and financial series, together with controlled simulations having exact posteriors. Across two fitted language models, language-derived probabilities outperform printed numerical confidence, recover held-out posteriors with valid uncertainty coverage, remain largely stable under paraphrase and respond appropriately to altered evidence. \textbf{Prompt engineering optimises a wording-dependent response; robust scientific use requires validated stability of application-relevant meaning.} The proposed map turns semantic uncertainty in generative systems into an identifiable and testable statistical measurement problem and, when its acceptance conditions hold, yields an auditable posterior estimate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑