世界杯竞技场:语言模型和深度研究代理在足球预测上的细粒度评估
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
浏览论文内容
中文总结 AI 辅助
介绍世界杯竞技场这一语言模型和深度研究代理的动态基准,用于足球预测。模型赛前接收证据或自行搜索信息进行多项预测,赛后与实际结果对比。通过多指标评估104场比赛和13个系统,展示不同模型表现及与基线对比情况,且新赛程可添加用于评估未来模型。
中文摘要 AI 辅助
开球前预测足球比赛需要的不仅仅是了解过去的结果:模型必须利用不断变化的信息并在答案可用之前做出明确预测。我们提出了世界杯竞技场,这是一个用于语言模型和深度研究代理的动态基准。2026年国际足联世界杯是其首次评估,相同流程可用于未来联赛和杯赛。每场比赛前,模型要么接收通用证据包,要么自行搜索信息,预测比赛结果、比分、可能的球员和事件、比赛统计数据以及比赛结果。比赛结束后,将这些预测与记录结果进行比较。我们报告了结果准确率、精确比分准确率、当预测比分接近但不准确时给予一定分数的比分线分数,以及其他预测任务的分数。在104场比赛和13个系统中,结果准确率相似的模型在详细预测上差异更明显。与博彩市场和球迷基线相比,最佳系统在结果和精确比分准确率上仅略有提高,但在比分线分数上有更明显的提高。新赛程开始时可添加,使该基准能够评估未来模型而无需使用已知结果。代码、提示、预测和评估脚本在该https网址开源。
英文摘要
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluation, and the same process can be reused for future leagues and cups. Before each match, a model either receives a common evidence package or searches for information itself. It predicts the result and score, likely players and events, match statistics, and the outcome of the competition. After the match, these predictions are compared with the recorded result. We report result accuracy, exact-score accuracy, and a scoreline score that gives some credit when a predicted score is close but not exact, together with scores for the other prediction tasks. Across systems, similar result accuracy can mask larger differences in detailed predictions. Four systems predicted champion Spain, and two of them also recovered the exact final pairing. Compared with betting-market and human-fan baselines, the best system shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline. New schedules can be added as they begin, allowing the benchmark to evaluate future models without using outcomes that are already known. Code, predictions and evaluation scripts will be publicly released.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- McGill University(麦吉尔大学)
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。