arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32885cs.AI

大型语言模型能否预测未来?预测市场的布赖尔评分分析

Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets

Yuanbo Li, Zekun Li, Xiaoyan cong

AI总结:

本研究通过布赖尔评分分析预测市场问题,发现模型升级未必提升概率估计准确性,强调需结合基线和问题构成解读预测分数。

AI中文摘要:

我们研究模型升级是否能提高预测市场问题的概率估计。我们的已解决市场预测(RMF)基准包含九个领域的3000个已解决二元问题,我们使用仅问题、零样本协议评估了六个Claude和Qwen模型变体。我们相对于经验基础率预测器评估布赖尔分数,检查其Murphy分解,并通过配对差异比较模型,结果按事件类别和相对于训练截止时间的时序进行分层。在报告的训练截止后层中,四个Claude模型实现了0.183-0.192的布赖尔分数,比基础率参考值提高了0.024-0.033。Qwen 32B没有显著优于该参考值,尽管其配对布赖尔分数比7B检查点低0.024。所评估的Claude版本和层级升级没有带来显著改进。模型内不同事件类别之间的差异超过了Claude变体之间的观察差异。这些结果表明,预测分数应结合简单概率基线和问题构成来解释:在此协议下,较新版本或较高模型层级并不一致地产生更准确的概率。

英文摘要:

We study whether model upgrades improve probability estimates for prediction-market questions. Our Resolved Market Forecasting (RMF) benchmark contains 3,000 resolved binary questions across nine domains, on which we evaluate six Claude and Qwen model variants using a question-only, zero-shot protocol. We assess Brier scores relative to an empirical base-rate predictor, examine their Murphy decomposition, and compare models through paired differences, with results stratified by event category and timing relative to training cutoffs. In the reported post-cutoff stratum, the four Claude models achieve Brier scores of 0.183-0.192, improving on the base-rate reference by 0.024-0.033. Qwen 32B does not significantly outperform that reference, although its paired Brier is 0.024 lower than that of the 7B checkpoint. The evaluated Claude version and tier upgrades yield no significant improvement. Within-model differences across event categories exceed the observed differences among Claude variants. These results show why forecasting scores should be interpreted alongside simple probability baselines and question composition: under this protocol, newer versions or higher model tiers do not consistently produce more accurate probabilities.

↑