arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

2026 AI世界杯:用于端到端足球锦标赛预测的大语言模型基准测试

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

Jonaid Shianifar, Iias Faiud

arXiv 2608.03416首次发表:更新:

AI 中文总结

本文通过AI世界杯基准测试,对10款大语言模型进行2026年FIFA世界杯端到端预测,发现淘汰赛表现决定排名,GPT-5.5 Thinking夺冠,相关成果已公开以支持后续研究。

AI 中文摘要

大语言模型(LLM)如今常被要求预测现实世界事件,但由于模型接收的信息不同、使用的工具不同、评估规则不同,对比往往存在困难。本文报告了已完成的AI世界杯基准测试,其中10个基于LLM的助手对2026年FIFA世界杯进行了赛前单次完整预测。所有提交内容均使用相同的锦标赛快照、提示词、JSON模式和评分流程。预测内容涵盖小组赛积分、小组排名、淘汰赛对阵表、最终名次、置信度值及简短解释。在全部104场比赛结束后,GPT-5.5 Thinking以744分位列第一,其次是GPT-5.5(717分)、Gemini(699分)和Qwen 3.7(687分)。GPT-5.5 Thinking也是唯一将西班牙选为冠军的模型,西班牙在决赛中1-0击败阿根廷。最终排名主要由淘汰赛表现决定:总分与淘汰赛积分呈强相关(r=0.986),但与小组赛比赛积分(r=0.055)、小组排名积分(r=-0.103)或淘汰赛之前的综合积分(r=-0.054)几乎无关。单场比赛准确率则产生了不同的排名:Claude Sonnet 4.6正确预测了最多的小组赛结果(63.89%),但总排名为第六。平均自我报告的置信度与结果准确率(r=-0.060)或总分(r=-0.067)也无关联。结果表明,预测完整锦标赛与逐场预测比赛所测试的内容不同,同时也显示基于对阵表的排行榜在多大程度上取决于评分设计。该基准测试材料、原始响应和评分代码已发布,以支持复现和未来扩展。

英文摘要

Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules. This paper reports the completed \emph{AI World Cup} benchmark, in which ten LLM-based assistants made a single pre-tournament forecast of the entire 2026 FIFA World Cup. Every submission used the same tournament snapshot, prompt, JSON schema, and scoring procedure. The forecasts covered group-stage scores, group rankings, the knockout bracket, final placings, confidence values, and short explanations. After all 104 matches had been played, GPT-5.5 Thinking finished first with 744 points, followed by GPT-5.5 with 717, Gemini with 699, and Qwen 3.7 with 687. GPT-5.5 Thinking was also the only model to select Spain, which defeated Argentina 1--0 in the final, as champion. The final ranking was driven mainly by knockout performance: total score was strongly correlated with knockout points ($r=0.986$), but showed little relationship with group-stage match points ($r=0.055$), group-standing points ($r=-0.103$), or their combined pre-knockout score ($r=-0.054$). Match-level accuracy produced a different ordering. Claude Sonnet 4.6 correctly predicted the largest number of group-stage outcomes (63.89\%) but placed sixth overall. Average self-reported confidence was also unrelated to either outcome accuracy ($r=-0.060$) or total score ($r=-0.067$). The results suggest that forecasting a complete tournament tests something different from predicting matches one at a time, while also showing how strongly a bracket-based leaderboard can depend on scoring design. The benchmark materials, raw responses, and scoring code are released to support replication and future extensions.

Comments18 pages, 8 figures, 7 tables. Project repository available in the paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑