arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NeutronGym:面向LLM智能体的物理分级中子仪器设计

NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

Lijie Ding, Changwoo Do

arXiv 2610.03631首次发表:更新:

发表机构

Oak Ridge National Laboratory(橡树岭国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

NeutronGym是首个中子仪器设计的可执行环境,通过物理分级阶梯和强化学习,使Qwen3-8B在留出任务上从11%提升至77%,超越未训练的更大模型,并揭示了奖励驱动的能力边界。

AI 中文摘要

设计科学仪器可以测试语言模型智能体是否真正理解物理而非仅仅回忆物理知识,前提是评分过程无可争议。我们引入了NeutronGym,据我们所知,这是首个用于中子仪器设计的可执行环境:智能体通过验证工具构建仪器,McStas对构建结果进行射线追踪模拟,一个分级阶梯对语法、运行、结构和科学性进行评分,不依赖LLM评判。程序化家族提供无限多个固定布局的实例,智能体必须设置其设计参数,并包含留出的参数区间;一个精选子集McStasBench增加了16个来自已发表仪器的任务,并配有记忆探测和沙箱。七个模型最多只能复现16个任务中的7个,没有一个能检索到参考方案,也没有一个达到改进目标。该环境还能用于训练。在其奖励上进行强化学习使Qwen3-8B在某个隐藏设计目标的家族留出实例上的表现从11%提升到77%(第二个种子为69%),超过了未训练的Qwen3-32B,并且该配方在另外三个门控家族上各用一个种子也成立。分析揭示了这一提升的本质。没有阶梯的部分学分,性能会下降60个百分点。仅凭奖励,训练后的模型在智能体的模拟预算下,只有获得闭式物理公式时才能达到经典优化器的水平(77%对81%,这一差距在此规模下不显著),而前沿模型仍能解决98-99%的任务。获得可信结果意味着要放弃四个无模型基线能解决的任务设计,我们发布了发现这些问题的探测工具。

英文摘要

Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with held-out parameter regimes; a curated slice, McStasBench, adds 16 tasks from published instruments behind memorization probes and a sandbox. Seven models reproduce at most 7 of the 16, none retrieves a reference, and none meets an improvement target. The environment also trains. Reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B, and the recipe holds, at one seed each, on three further gated families. The analysis says what that gain is. Without the ladder's partial credit it collapses by 60 points. From reward alone the trained model reaches what a classical optimizer reaches, at the agent's simulation budget, only when handed the closed-form physics (77% against 81%, a gap that does not separate at this size), while frontier models still solve 98-99%. Getting a trustworthy result meant failing four task designs that no-model baselines could solve, and we release the probes that found them.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑