arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12101cs.AI

语言模型与先验的能力门控池化用于事件预测

Competence-Gated Pooling of Language Models and Priors for Event Forecasting

Aditi Tiwari, Aashrith Bandaru, Heng Ji

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出能力门控机制,通过估计领域级源权重并收缩校准,在事件预测中显著提升外部基线性能,并支持基于边际价值的模型选择性使用与弃权决策。

中文摘要 AI 辅助

在混合预测中,语言模型通常是多种可用信号之一。系统可能已有市场、群体或统计预测,必须决定该模型是增加有用信息还是应被忽略。因此,相关目标不是模型独立准确性,而是相对能力,定义为模型在现有外部预测之外的边际价值。在布里尔损失下,我们刻画了模型分歧何时能改善外部预测,并推导了使用领域特定而非全局池化权重的收益。然后,我们引入一个能力门,从已解决的结果中估计领域级源权重,将不确定估计向全局权重收缩,并重新校准池化预测。在2,357个已解决的二元问题和五个语言模型上,该门将主要外部基线从0.0771改善至0.0762的布里尔分数(注:原文为0.0732,此处按原文保留),并显著优于全局预测组合。在泄漏控制和针对池化结构化集的防泄漏时间序列先验下,增益仍然显著,且在FRED上有独立证据。相比之下,该门在官方ForecastBench市场子集上无显著改进,在该子集上它主要服从市场。在四个Qwen模型中,口头置信度不能可靠地识别模型何时优于外部预测,而结果估计的能力支持更好的弃权(不执行)决策。这些结果为基于测量边际价值的选择性模型使用提供了实用方法。

英文摘要

In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast. Under Brier loss, we characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights. We then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, the gate improves the main external baseline from 0.0771 to 0.0732 Brier and significantly outperforms global forecast combinations. The gain remains significant under leakage controls and against a leakage-safe time-series prior on the pooled structured set, with separate evidence on FRED. In contrast, the gate gives no significant improvement on the official ForecastBench market subset, where it largely defers to the market. Across four Qwen models, verbal confidence does not reliably identify when the model outperforms the external forecast, while outcome-estimated competence supports better abstention decisions. These results provide a practical approach for selective model use based on measured marginal value.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑