arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从得分矩阵到足球感知的比赛状态模拟:用于精确得分重排序的可审计大语言模型(LLM)框架

From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking

Shaopeng Liang

arXiv 2608.05030首次发表:更新:

AI 中文总结

本研究提出可审计混合架构,结合动态统计模型与LLM实现足球得分重排序,经迭代优化后在英超150场比赛时序测试中提升了精确得分准确率与候选覆盖度,明确了该方法的改进及局限。

AI 中文摘要

足球得分预测兼具扎实的统计核心与复杂的语境挑战:动态泊松族模型可估算球队实力、预期进球数及连贯得分概率,但无法直接理解球员角色、战术对阵、动机或首粒进球如何改变球队行为。大语言模型(LLM)可推理此类概念,却并非校准后的概率引擎。本文通过可审计信息框架将二者结合,记录了四个迭代版本:V1为动态得分驱动的Dixon-Coles基线模型;V2将LLM语境评分映射回预期进球参数;V3用冻结得分候选集上的逐球模拟替代标量修正;V4新增共享首破局与进球后连锁判断、时间感知停止机制及确定性尾部候选。该框架定义输入语义、提供赛前证据,并约束LLM进入可检查的推理路径。在2025-26赛季英格兰超级联赛前150场比赛的时序重放测试中,V1的Top-1精确得分准确率为10.0%、Top-3为26.7%;V3分别达12.0%、30.0%;V4则达14.7%、30.7%,且候选覆盖度从77.3%提升至84.7%,但新增尾部候选未产生Top-3精确命中。V1原生1X2分布的argmax准确率为53.3%,对数损失0.9878,Brier分数0.5870,排序概率分数0.2095。这些结果为探索性的:开发切片并非未触及的基准,且时序输入隔离无法排除封闭LLM中的结果记忆。本研究的贡献为可审计混合架构、清晰的设计演进,以及揭示足球感知模拟在得分选择上的改进与局限的负面结果。

英文摘要

Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour. Large language models (LLMs) can reason about such concepts, yet are not calibrated probability engines. We combine both components through an auditable information harness. This paper documents four iterations: V1, a dynamic score-driven Dixon-Coles baseline; V2, which maps LLM contextual ratings back into expected-goal parameters; V3, which replaces scalar correction with goal-by-goal simulations over a frozen score-candidate set; and V4, which adds shared first-breakthrough and post-goal cascade judgments, time-aware stopping, and deterministic tail candidates. The harness defines input semantics, supplies pre-match evidence, and constrains the LLM to an inspectable reasoning route. On a chronological replay of the first 150 matches of the 2025-26 English Premier League, V1 achieved 10.0% Top-1 and 26.7% Top-3 exact-score accuracy. V3 reached 12.0% and 30.0%, while V4 reached 14.7% and 30.7%. V4 increased candidate coverage from 77.3% to 84.7%, although no added tail candidate became a Top-3 exact hit. V1's native 1X2 distribution achieved 53.3% argmax accuracy, 0.9878 log loss, 0.5870 Brier score, and 0.2095 ranked probability score. These results are exploratory: the development slice is not an untouched benchmark, and temporal input isolation cannot exclude outcome memory in a closed LLM. The contribution is an auditable hybrid architecture, a clear design evolution, and negative findings showing where football-aware simulation does and does not improve score selection.

Comments9 pages, 1 figure, 5 tables. Interim chronological benchmark on the first 150 matches of the 2025-26 English Premier League

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑