对抗性快速变化的真实世界领域:用于基准测试AI科学家能力的测试床
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
浏览论文内容
中文总结 AI 辅助
该研究提出以F1、MTG等对抗性快速变化的真实世界领域为测试床,发现AI科学家的核心能力差距是想法过滤、优先级排序而非生成。
中文摘要 AI 辅助
对AI科学家生成新颖想法的能力进行基准测试向来极为困难。该领域现有的基准在评估科学推理和研究复现方面已取得进展,但往往依赖合成任务或回顾性目标,可能因先前接触而产生混淆。我们假设,复杂、对抗性且快速变化的真实世界领域(其中专家从业者独立生成可观测输出)可提供实用解决方案,以填补这一空白并评估AI科学家所需的能力,包括推理、新颖性和假设构建。我们在两个结构不同的领域实例化该框架:一级方程式(F1)领域,模型围绕2026赛季的赛车设计概念构思,而真实的季前创新构成了真值;万智牌(MTG)领域,模型从近期更新的卡牌池提出牌组,并对照19个职业巡回赛(PT)牌组进行评估。在两个领域中,模型均生成了看似合理的输出,但很少与真实世界专家的解决方案对齐。在F1领域,最佳模型GPT-5.2在多次运行中提出166个想法,匹配了40项真实创新中的10项;在MTG领域,Gemini 3 Flash生成的最佳牌组从第三名PT牌组中找回了7张新系列卡牌中的5张,且在全部108个牌组中,模型选择频率最高的卡牌也是PT牌组最广泛采用的卡牌(斯皮尔曼ρ=0.74,p=0.0003)。这些结果表明,AI科学家的关键能力差距不在于想法生成,而在于过滤、优先级排序和连贯的新颖性。
英文摘要
Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. We hypothesize that complex, adversarial, fast-moving real-world domains where expert practitioners independently generate observable outputs can provide a practical solution to fill this gap and evaluate the capabilities needed for AI scientists, including reasoning, novelty, and hypothesis formulation. We instantiate this framework in two structurally different domains, Formula 1 (F1), where models ideate around car design concepts for the 2026 season, and real pre-season innovations provide a ground truth, and Magic: The Gathering (MTG), where models propose decks from a recently updated card pool and are evaluated against 19 Pro Tour (PT) decklists. Across both domains, models produce plausible outputs, but few align with real-world expert solutions. In F1, the best model, GPT-5.2 matched 10 of 40 real innovations with 166 ideas proposed across runs. In MTG, the best deck from Gemini 3 Flash recovered 5 of 7 new-set cards from the third-place PT deck, and across all 108 decks, the cards models selected most often were also the cards most widely adopted by PT decks (Spearman $ρ= 0.74$, $p = 0.0003$). These results suggest that a key capability gap for AI scientists is not idea generation, but filtering, prioritization, and coherent novelty.