arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18423cs.AI

FM-Bench:用于多智能体竞争的长周期管理基准

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

Tianyou Wang, Chongyang Gao, Kezhen Chen, Dong Chen, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出FM-Bench基准,通过足球管理任务评估15个前沿LLM智能体的长周期决策能力,发现模型管理行为而非计算能力决定得分,排名与模型规模、价格或供应商无关。

中文摘要 AI 辅助

语言模型智能体现在能够可靠地执行有界任务,但它们能否在长周期中维持有效决策(其中行动具有累积后果且环境会对其选择做出响应),在很大程度上仍未得到测量。FM-Bench(足球管理基准)正是用于测量这一能力。一个大语言模型(LLM)智能体通过26种工具,在约340至400个决策节点中运营一家足球俱乐部,时长为20个游戏内年份。它与所有竞争对手使用相同预算选拔阵容、交易球员、谈判合同、投资设施与青训、设置首发阵容,并需应对可能解雇它的董事会;同时,一个确定性引擎每年累积数据,最终生成最终得分,全程无需LLM评判或人工评分者。单人赛道将15个前沿模型分别置于固定的脚本化世界中,竞技场赛道则将相同模型与一个脚本化锚点置于同一个共享的20年世界中;据我们所知,这是首次开展此类规模的头对头评估。我们测量了得分背后的六种行为能力。在3个随机种子下,所有15个模型均完成了全部周期,而盲脚本化基线在其多数周期中失败;claude-fable-5在单人赛道和竞技场赛道的平均得分中位居榜首,不过冠军头衔在10个模型间轮换。模型的排名与规模、价格或供应商均无关联;排名仅在周期后期才稳定下来,且首个参赛的人类选手仅处于模型榜单的末尾。区分模型的关键在于管理行为而非计算能力:得分更高的模型会在周期末期减少慢回报投资、保持现金用于投资而非闲置、在截止日期前很久就开启续约,而token消耗与得分无关。没有模型能从数百次被拒绝的出价中学习到市场的隐藏价格,且自管理记忆存在两种相反的失效模式:要么是仅增长的档案,要么是每个赛季都重写的计划。代码可在此https URL获取。

英文摘要

Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater. The solo track plays each of 15 frontier models against a frozen scripted world, and the Arena places the same models plus a scripted anchor in one shared 20-year world; to our knowledge, the first head-to-head evaluation at this scale. We measure six behavioral capabilities behind the score. Across three seeds, all 15 models complete every horizon while the blind scripted baselines die out in most of theirs, and claude-fable-5 tops the solo board on mean score and the Arena, where the title nonetheless rotates among ten models. Neither scale, price, nor vendor predicts the order; the order settles only late in the horizon, and the best first-play human lands only at the bottom of the model board. What separates the models is managerial behavior rather than computation. Higher-scoring models reduce slow-payoff investment near the end, keep cash invested rather than idle, and open renewals well before the deadline, while token spend predicts nothing. No model learns the market's hidden prices from hundreds of rejected bids, and self-managed memory fails in two opposite modes: an archive that only grows or a plan rewritten every season. Code is available at https://github.com/Analogy-AI/fm-bench.

发表机构

  • AnalogyAI

机构由 AI 辅助整理,请以论文原文为准。

↑