arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FromPitch2Board:长时程足球管理中的LLM智能体基准测试

FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management

Peiyu Zang

arXiv 2609.34710首次发表:更新:

发表机构

Beijing Normal University; Peking University(北京师范大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出FromPitch2Board确定性足球管理基准,通过受控比较分离模型、脚手架、责任范围等因素,揭示GPT-5.6在责任扩展下被动性剧增及长时程行为变化。

AI 中文摘要

长时程智能体基准通常报告智能体能够推进多远,但无法识别其性能究竟源自基础模型、脚手架、责任范围、比赛控制粒度还是时程。我们提出FromPitch2Board,一个确定性的足球管理基准,通过在单一模拟器上进行受控比较来研究五个可配置因素,使用配对种子和冻结校准。我们评估了四个基础模型和四个智能体脚手架。在模型轨道中,教练得分Z分数跨度0.19,而经理得分Z分数跨度0.68,其中GPT-5.6在责任扩展下被动性急剧上升。其责任阶梯从46.1分升至58.1分(含招募),但在全面管理下降至46.8分,将回归定位在最终责任边界。跨越该边界,其跳过决策率从1.1%升至57.9%。在跨越每个脚手架的Flash-Pro配对中,相对于固定无状态脚手架,脚手架选择使经理得分Z分数变化高达0.48。3Y队列在一年级和三年级之间显示出平均排名的方向性反转,而选定的Claude Code+Pro配置在三年级达到峰值并保持在该峰值之下,表明责任范围和时程暴露了单一标题分数所掩盖的行为变化。

英文摘要

Long-horizon agent benchmarks typically report how far an agent progresses, but do not identify whether its performance comes from the foundation model, scaffold, responsibility scope, match-control granularity, or horizon. We introduce FromPitch2Board, a deterministic football-management benchmark that studies five configurable factors through controlled comparisons on a single simulator, using paired seeds and a frozen calibration. We evaluate four foundation models and four agent scaffolds. In the Model Track, Coach points Z-scores span 0.19, while Manager points Z-scores span 0.68, with GPT-5.6 showing a sharp rise in passivity under responsibility expansion. Its responsibility ladder rises from 46.1 to 58.1 points with recruitment, then falls to 46.8 under full management, localizing the regression to the final responsibility boundary. Across that boundary, its skipped-decision rate rises from 1.1% to 57.9%. Within the Flash-Pro pair crossed across every scaffold, scaffold choice changes Manager points Z-scores by up to 0.48 relative to the fixed stateless scaffold. The 3Y cohort shows a directional reversal in mean ranking between years one and three, while a selected Claude Code+Pro configuration peaks in year three and remains below that peak, showing that responsibility scope and horizon expose behavior changes that a single headline score conceals.

Comments28 pages, 4 figures. Code: https://github.com/factnn/FromPitch2Board

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑