Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
Running the Gauntlet: 重新评估智能体在陌生环境中的能力
Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel, Michal Zakrzewski, Sebastian Montagna, Damian Rynczak, Shreyansh Padarha, Kumail Alhamoud, Zihao Fu, William Lugoloobi, Kai Rawal, Hanna Yershova, Xander Davies, Taras Rumezhak, Guohao Li, Fazl Barez, Baoyuan Wu, Arkadiusz Drohomirecki, Yarin Gal, Chris Russell, Christopher Summerfield, Adam Mahdi, Volodymyr Karpiv, Philip Torr, Adel Bibi
机构
*
University of Oxford(牛津大学)
;
SoftServe
;
Massachusetts Institute of Technology(麻省理工学院)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
UK AI Security Institute(英国人工智能安全研究所)
;
Ukrainian Catholic University(乌克兰天主教大学)
When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations
当故事演变时:在开放世界模拟中针对智能体架构对大语言模型故事讲述能力进行基准测试
Yuqi Chen, Sixuan Li, Yunfeng Cai, Xueai Li, Ka Man Yan, Ying Li
机构
*
The University of Hong Kong(香港大学)
;
Peking University(北京大学)
;
Tsinghua University(清华大学)
;
Beijing Institute of Mathematical Sciences and Applications (BIMSA)(北京数学科学与应用研究院)
Comments25 pages, 4 figures, 16 tables, 6 appendices. Code, task suite, released per-run verdicts, and a one-command reproduction of every reported number: https://github.com/shivenkk/agentrelbench
The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
红皇后哥德尔机:共同进化的智能体及其评估者
Alex Iacob, Andrej Jovanović, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, Lorenzo Sani, Niccolò Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane
Comments15 pages, 3 figures. Reliability protocol and library (cua_reliability), completion verifier, and leakage-free held-out plus bootstrap evaluation harness are open-source