AI 中文总结
本研究构建世界杯竞技场基准,在2026世界杯期间让6个前沿LLMs实时预测,发现其平均准确率63.9%,无明显差异,将相关数据和代码作为基准发布。
AI 中文摘要
用于评估大语言模型预测能力的基准几乎都是回顾性的:事件已经发生,答案存在于网络某处,评估必须防御记忆问题。我们报告一种相反的设计:在2026年国际足联世界杯的39天里,6个前沿大语言模型——均具备扩展思考能力和原生服务器端网页搜索功能——在每场开球前,一次一场比赛,被要求填写包含所有104场比赛的7项市场预测卡,外加12个小组冠军和赛前夺冠赔率池;提问时不存在任何答案,因此该评估从构造上而非通过过滤实现无泄露,冻结的存档包含4494个已评分的预测。该赛事确立了6个系统的一系列共同行为:在比赛结果预测上,它们的平均准确率为63.9%,与支持博彩公司的热门队伍相当——而它们通常确实这么做;它们彼此间的一致性远高于正确率,因此多数投票没有增益;它们对平局和进球的预测不足,且将比分选择集中在单一典型结果上;准确率与比赛的悬殊程度相关,而非与对比赛的了解程度相关:在信息最丰富的最接近的比赛中,准确率崩溃,而关于赛事整体的问题回答良好。在该任务上,当前一代前沿系统没有明显差异:排名在整个过程中上下波动,顶部和底部位置保持稳定,中间的排名变动,且差距始终很小。我们将赛事简报档案、赛程和官方结果连同评分代码一起作为基准发布。
英文摘要
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the 2026 FIFA World Cup, six frontier LLMs -- all with extended thinking and native server-side web search -- were asked before every kickoff, one match at a time, to fill in a seven-market prediction card for all 104 matches, plus 12 group winners and a pre-tournament outright pool; no answer existed when the question was asked, so the evaluation is leakage-free by construction rather than by filtering, and the frozen archive holds 4,494 scored predictions. What the tournament establishes is a set of behaviours the six systems share. On match outcome they average 63.9%, level with backing the bookmaker's favourite -- which is in fact what they usually do. They agree with one another far more often than they are right, so a majority vote adds nothing. They under-commit to draws and to goals, and crowd their scoreline picks onto a single prototypical result. Accuracy tracks how lopsided a fixture is rather than how much is known about it: it collapses in the closest ties, where the dossiers are richest, while questions about the tournament as a whole are answered well. On this task the current generation of frontier systems is not sharply differentiated: the standings hold up at the top and the bottom across the run and churn in the middle, and the margins stay narrow throughout. The briefing dossiers, fixtures and official results are released as a benchmark, together with the scoring code.
CommentsProject page: https://co-minder.github.io/worldcup2026