arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Witness:交互式谜题环境中的发现、解读与顿悟

Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments

Guanghan Ning, Ping Liu, Linyi Li, Huangjie Zheng, Arjun Neervannan, Huu Nguyen, Michael Sklar, Deniz Zorlu, Nicolai Ouporov

arXiv 2609.32208首次发表:更新:

AI 中文总结

本文提出WITNESS交互式谜题环境,研究语言模型在规则发现上的局限,并通过强化学习提升模型在未见规则上的表现,揭示规则获取是前沿模型的主要瓶颈。

AI 中文摘要

自动化科学需要能够通过与环境交互来弄清陌生环境规则的智能体。交互式规则发现谜题为此能力的研究提供了一个受控环境:智能体通过实验推断隐藏规则,并利用推断出的规则达到既定目标。我们探究了当前语言模型在这些谜题上的局限,以及强化学习(RL)是否能在训练中未出现的规则上提升性能。为研究这两点,我们引入了WITNESS,一个基于2D网格的谜题环境,具有真实ASCII观测和对规则的受控访问。一个智能体流水线为WitnessGym(RL训练套件)和WitnessBench(包含公开验证集和私有测试集游戏)生成游戏。验证集分别测试训练规则原语的新组合和训练中未出现的原语。在共享测试框架下,18个前沿专有和开源权重模型中最好的仅解决了私有测试关卡槽位的24%,且得分对观测接口和智能体配置敏感。提供真实规则将Opus-5的验证RHAE-L5(前五关相对人类行动效率)从59.9提升至97.8,而一个27B开源权重模型仅提升2.1点,即使提供规则仍受限。在WitnessGym上进行RL将27B模型的私有测试RHAE-L5从2.1提升至5.4,并在四个外部发现基准上平均提升4.1点。这些结果共同表明,规则获取是Opus-5等前沿模型的主要难点,而较小模型在基于规则的执行上更为挣扎,并表明在隐藏规则谜题上的RL可迁移到更广泛的规则和训练之外的真实世界任务。基准可在以下网址获取:this https URL

英文摘要

Automated science needs agents that can work out the rules of an unfamiliar environment by interacting with it. Interactive rule-discovery puzzles offer a controlled setting for studying this ability: an agent infers hidden rules through experimentation and uses what it has inferred to reach a stated goal. We ask what limits current language models on these puzzles and whether reinforcement learning (RL) improves performance on rules held out from training. To study both, we introduce WITNESS, a 2D grid-based puzzle environment with ground-truth ASCII observations and controlled access to rules. An agentic pipeline generates games for WitnessGym, the RL training suite, and WitnessBench, comprising public validation and private test games. The validation set separately tests new compositions of trained rule primitives and primitives absent from training. Under a shared harness, the best of 18 frontier proprietary and open-weight models solves only 24\% of private test level slots, with scores sensitive to the observation interface and agent configuration. Providing ground-truth rules raises Opus-5's validation RHAE-L5 (relative human action efficiency over the first five levels) from 59.9 to 97.8, whereas a 27B open-weight model gains only 2.1 points and remains limited even with the rules provided. RL on WitnessGym raises the 27B model's private test RHAE-L5 from 2.1 to 5.4 and yields a mean gain of 4.1 points on four external discovery benchmarks. Together, these results point to rule acquisition as a major difficulty for frontier models like Opus-5 while smaller models further struggle on rule-based execution, and indicate that RL on hidden-rule puzzles transfers to broader rules and real-world tasks beyond training. Benchmark is available at: https://witnessbench.ai

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑