arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ExplorationBench:在可验证的外星世界中衡量AI系统的探索能力

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan

arXiv 2609.30199首次发表:更新:

发表机构

Fudan University; Hunyuan Team Tencent; Tsinghua University(复旦大学; 腾讯混元团队; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ExplorationBench通过可执行规则与冲突知识的外星世界沙盒,将科学探索评估转化为可验证任务,实验表明最强AI能获取新规则但性能波动且探索可能停滞。

AI 中文摘要

科学发现始于已知问题终结之处。在那里,AI系统必须进行探索:提出假设、设计实验,并基于结果进行迭代。然而,评估这种能力是困难的:(1)如何验证一个真正的新假设是否成立,(2)如何确定系统是通过探索发现它,还是仅仅从预训练数据中回忆了相关知识。为此,我们引入了ExplorationBench,它将评估科学探索这一棘手问题转化为一个具体且可处理的框架,该框架建立在可验证的外星世界之上:这些世界的规则是可执行的,因此每个答案都可以被精确检查,并且它们与熟悉的知识相冲突,因此仅靠回忆无法解决任务。该基准包含两个沙盒:AlienCode(31个发现目标,70个任务)和AlienLogic(24个发现目标,70个任务)。每个沙盒提供一份有缺陷的手册、特定于任务的环境反馈,以及专门的工具调用模式。系统利用这些资源探索沙盒,然后解决保留任务。我们评估了10个AI系统,发现最强的系统能够获取并应用不熟悉的规则,而性能在不同轨迹间差异很大,持续探索可能会停滞或逆转早期的进展。ExplorationBench代表着朝着AI系统能够在未知环境中通过探索获取并应用真正新知识迈出的一步。

英文摘要

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark contains two sandboxes, AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbox provides a flawed manual, task-specific environmental feedback, and a dedicated tool-call schema. Systems use these resources to explore the sandbox, then solve held-out tasks. We evaluate 10 AI systems and find that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can stall or reverse earlier gains. ExplorationBench represents a step towards AI systems that can acquire and apply genuinely new knowledge through exploration in unknown environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑