AI 中文总结
该研究提出前沿自动实验室(Frontier Autolab),一个跨越1990至2040年九个技术时代的多智能体LLM企业长时程测试平台,发现远见-承诺差距及组织记忆对长期行为的影响,并揭示评估中的时间泄漏问题。
AI 中文摘要
多智能体LLM系统日益被构建为类似组织的结构,配备角色、批评者和共享记忆,然而它们却是在持续数分钟的任务上进行评估。我们提出疑问:当这样的组织所立足的基础不断变化时,它将如何表现?前沿自动实验室(Frontier Autolab)是一个长时程测试平台,其中一个模拟企业由十六个角色人设和一支专职红队(Red Team)配音,必须在从1990年到2040年的九个技术时代中重新自我奠基。每个时代都有时间门控:企业根据一份注明日期的简报做出决策,随后一位历史学家-法官揭示实际发生的情况,并在五个维度的评分标准上对决策进行评分,而经验教训则进入一个持久化的剧本手册(Playbook)。六个时代依据历史进行评分,一个时代针对实时市场进行评分,还有两个时代是开放式预测。在四条轨迹(36个时代决策,180个子评分)中,我们发现了一致的远见-承诺差距:在所有24个历史评分时代中,法官对企业识别即将到来的转变的评分高于其对建设地点的选择(在10分制上平均差距为1.9分),因为董事会选择了其现有资产能够触及的层级。组织设计塑造了长期特征。一支配备数值终止门(kill gates)的红队产生了五十年的门控试点且没有产品,而评分标准将该企业评为最高;记忆存储了市场结构经验教训的企业在每个时代都进行了转向,而记忆仅存储验证程序的企业则始终固守一种方法。我们还展示了为何此类结果难以令人信服。在每次运行中,分数跨时代上升,而法官自身的后见之明子评分却下降(运行内相关系数r = -0.58),因此表面上的学习与对历史的回忆相混淆,我们进一步将扭曲追溯到自我评判、简报选择和分数聚合。我们发布了所有记录和一个API接口,并指定了虚构和截止日期后的时代,以将该测试平台转变为基准测试。
英文摘要
Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge's own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.
Comments17 pages, 6 figures, 9 tables. Code and data: https://github.com/LoopGlitch26/Frontier-Autolab