arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17765cs.LGcs.AIcs.DB

2026年国际足联世界杯作为语言模型预测代理的无污染基准:四个模型、一个博彩公司和104场比赛

FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

Jiacheng Ding, Cong Guo, Jason Xu

首次发表
浏览论文内容

中文总结 AI 辅助

介绍WC2026-Agents基准和数据集,用于评估大语言模型作为世界杯预测代理。四个前沿模型对104场比赛运行相同循环,与博彩市场数据配对。揭示模型预测虽有相同首选,但在决策质量等方面差异大,可衡量校准等方面。

中文摘要 AI 辅助

我们引入了WC2026-Agents,这是一个用于评估大语言模型(LLMs)作为真实未来事件自主预测代理的基准和数据集。对于2026年国际足联世界杯的104场比赛中的每一场,四个前沿模型——Claude Opus 4.8、ChatGPT(GPT-5.5,高推理)、Gemini 3.1 Pro和Grok(专家模式)——运行相同的搜索-行动-反思循环:使用网络工具收集证据,确定1X2(A队胜/平/B队胜)分布和虚拟100美元赌注,比赛结束后仅根据最终比分进行反思。由于每场比赛在模型训练截止日期后开始,该基准在构建上是无污染的。至关重要的是,我们将这四个代理与从相同信息环境——赛前博彩市场——中抽取的第五个竞争对手配对,收集每场比赛的1X2赔率,给出一个经济上有依据的基线,让我们不仅能评估代理预测了什么,还能评估它如何处理资金。该版本包含416个预测和414个带有逐字推理、地面真相(包括点球大战)、赔率和可重现评估套件的反思。一个参考评估揭示了原始准确性掩盖的结果:这四个代理在92%的比赛中给出相同的首选,没有一个超过市场的布里尔分数;事实上,对市场热门进行简单的等额下注比这四个代理都赚得多。然而,作为决策者,这些代理有很大差异:投注投资回报率从-18%到+10%不等,对所有四个代理来说,背离市场都是无利可图的,引用市场的预测份额从12%到100%不等,对错误选择的自我报告错误率从36%到86%不等。因此,该基准衡量了校准、决策质量和自我认知——即使前沿模型的预测没有差异,它们在这些方面也存在差异。数据和代码:此https URL

英文摘要

We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events. For every one of the 104 matches of the 2026 FIFA World Cup, four frontier models -- Claude Opus 4.8, ChatGPT (GPT-5.5, high reasoning), Gemini 3.1 Pro, and Grok (Expert Mode) -- ran an identical search-act-reflect loop: gather evidence with a web tool, commit to a 1X2 (team-A win / draw / team-B win) distribution and a virtual 100-USD bet, and, after the match, reflect given only the final score. Because every match kicked off after the models' training cutoffs, the benchmark is contamination-free by construction. Crucially, we pair the four agents with a fifth competitor drawn from the same information environment -- the pre-match betting market -- collected as per-match 1X2 odds, giving an economically grounded baseline and letting us score not just what an agent predicts but what it does with money. The release contains 416 forecasts and 414 reflections with verbatim reasoning, ground truth (including penalty shootouts), odds, and a reproducible evaluation suite. A reference evaluation surfaces findings that raw accuracy hides: the four agents issue an identical top pick in 92% of matches and none beats the market's Brier score; indeed, a naive flat stake on the market favorite out-earns all four agents. Yet the agents diverge sharply as decision-makers: betting return-on-investment ranges from -18% to +10%, fading the market is unprofitable for all four, the share of forecasts that cite the market ranges from 12% to 100%, and self-reported error rates on wrong picks range from 36% to 86%. The benchmark thus measures calibration, decision quality, and self-knowledge -- axes on which frontier models differ even when their predictions do not. Data and code: https://github.com/graphuofm/FIFA2026LLM

发表机构

  • University of Memphis(孟菲斯大学)
  • QuantaInsight(量子洞察)

机构由 AI 辅助整理,请以论文原文为准。

↑