发表机构
MCML & LMU Munich; Paderborn University; University of Cologne(慕尼黑机器学习中心及慕尼黑大学; 帕德博恩大学; 科隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在评估大语言模型预测现实世界体育赛事的能力,核心方法是引入LLM - 足球竞技场这一前瞻性实时基准,通过对2026年世界杯的评估,详细分析模型性能各维度表现,为LLMs预测性能提供新证据及开源平台。
AI 中文摘要
大语言模型(LLMs)越来越多地支持对不确定未来事件的决策,但评估其预测现实世界结果的能力仍然困难。现有基准通常是静态和回顾性的,无法测试LLMs如何在不确定性下综合信息预测未来事件。我们引入LLM - 足球竞技场,这是一个前瞻性实时基准,用于评估LLMs在结果未知前对现实世界体育赛事的预测能力。它提供了前瞻性实时基准协议、公共开源平台以及因子基准设计和赛事相关问题。通过对2026年国际足联世界杯的大规模评估展示了该基准,对模型在信息获取、提示策略和预测范围等方面的性能进行了详细分析,为最先进的LLMs的预测性能提供了新证据。例如,有网络访问权限的LLMs比没有的表现稍好。总体而言,LLM - 足球竞技场为未解决事件的前瞻性基准测试提供了灵活的开源平台,将持续更新并可直接应用于未来赛事。
英文摘要
Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLMs generated forecasts for all 104 matches and 15 tournament-related questions. We provide a detailed analysis of model performance across information access, prompting strategy, and forecast horizon. As a result, LLM-SoccerArena provides new evidence about the forecasting performance of state-of-the-art LLMs. For example, LLMs with web access outperform those without, but only by a small margin (i.e., a 0.023 improvement in Brier score). Overall, LLM-SoccerArena provides a flexible, open-source platform for prospective benchmarking of unresolved events. LLM-SoccerArena will be continuously updated, and can be directly applied to future national and international tournaments and league competitions.