arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HEPLocalAgent 1.0:在本地计算机上基于纯文本请求运行对撞机模拟

HEPLocalAgent 1.0: Running Collider Simulations from Plain-Language Requests on Your Own Computer

Aadarsh Singh, Sudhir K. Vempati

arXiv 2608.28244首次发表:更新:

AI 中文总结

该研究提出开源本地界面HEPLocalAgent 1.0,可从自然语言请求构建对撞机模拟工作流,经测试其在47个案例中部分指标表现良好,需专家批准而非自主运行。

AI 中文摘要

我们提出了HEPLocalAgent,这是一个开源本地界面,可从自然语言请求构建一类有界的对撞机模拟工作流。本地部署的语言模型会生成带类型的工作流表示,确定性软件随后会恢复用户声明的已识别量、构建HEP工具输入、验证受支持的工作流,并在执行前展示工件以供批准。在对47个可评估模型请求案例的相同响应比较中,首次结构化提议给出了7个未修改的、满足外部基准评分器要求的工件,而经过完整确定性流水线后得到19个;在基准定义的固定表示归一化下,这两个数字分别为11和43,改进方向不变。两种视图的差异源于发布的构建器与基准评分器在三个簿记约定上存在分歧,即启动形式、两条固定控制行和输出目录名称,而非物理内容。7个获批工作流中有4个在托管本地软件栈上运行完成,重复运行的截面一致。在单独的挑战集中,96个有问题的请求中有57个在请求的部分内容被删除、默认设置或重新解释后仍进入了批准阶段。在批准关口前,可执行工件中未保留任何经测试的不安全有效载荷,但这并未建立操作系统级别的隔离。确定性后端支持MadGraph、Pythia8、Delphes和受限的MadAnalysis 5计划,在测试示例中未展示可靠的自然语言路由到MadAnalysis阶段的能力。因此,1.0版本应被视为一个可检查的、带验证关口的工作流构建器,需要专家批准,而非自主或科学自验证智能体。

英文摘要

We present HEPLocalAgent, an open source local interface that builds a bounded class of collider simulation workflows from natural language requests. A locally served language model proposes a typed workflow representation, and deterministic software then restores recognized user stated quantities, builds the HEP tool inputs, validates the supported workflow, and presents the artifacts for approval before execution. In a same response comparison on 47 evaluable model request cases, the first structured proposal gave 7 unmodified artifacts satisfying the external benchmark scorer, against 19 after the full deterministic pipeline. Under the fixed representation normalization defined by the benchmark, the counts were 11 and 43. The direction of improvement is unchanged. The gap between the two views arises because the released builder and the benchmark scorers disagree on three bookkeeping conventions, namely launch form, two fixed control lines, and the output directory name, not on physics content. Four of seven approved workflows ran to completion on the managed local software stack, with cross sections consistent between repeats. In a separate challenge set, 57 of 96 problematic requests still reached the approval stage after part of the request was dropped, defaulted, or reinterpreted. No tested unsafe payload was retained in an executable artifact before the approval gate, but this does not establish operating system level containment. The deterministic backend supports MadGraph, Pythia8, Delphes, and a restricted MadAnalysis 5 plan. Reliable natural language routing to the MadAnalysis stage was not demonstrated in the tested examples. Version 1 should therefore be seen as an inspectable, validation gated workflow constructor requiring expert approval rather than an autonomous or scientifically self validating agent.

Comments24 pages, 9 tables, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑