arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HEPToolBench 1.2:测试语言模型驱动粒子物理软件的可靠程度

HEPToolBench 1.2: Testing How Reliably Language Models Can Drive Particle Physics Software

Aadarsh Singh, Sudhir K. Vempati

arXiv 2608.28232首次发表:更新:

AI 中文总结

该研究推出HEPToolBench 1.2基准,评估42个语言模型驱动粒子物理软件的性能,发现结构化接口可大幅提升任务通过数,且该提升并非源于提取规则差异。

AI 中文摘要

科学家越来越希望通过自然语言请求驱动研究软件,但只有当流畅的输出能转化为正确的机器可读工件时才有实际意义。我们推出HEPToolBench,这是一个包含28个对撞机模拟任务的基准,由确定性的任务特定评分器评分,还附带一个包含3个任务的结构化调试扩展。我们评估了42个部署实例,范围从本地部署的小型开放权重模型到托管的前沿系统。核心实验比较了两种方式:直接生成原生HEP工具语法,以及通过模式介导的接口,模型返回类型化表示,由确定性软件进行序列化。在5个匹配请求中,平均得分从0.418上升到0.902,任务通过数从21/210提升至159/210,42个部署实例中有41个表现出提升;11个部署实例(其中7个是本地部署的开放权重模型)在原生语法下的5个任务通过数为0,在结构化接口下提升至5/5。数值使用修正后的v1.2.1契约,其中两个原生评分器此前强制执行其提示中未包含的操作约定,现按提示一致性对所有群组重新评分。仍存在一个不对称性:原生评分器逐字记录响应并拒绝Markdown围栏回复,而结构化评分器可从此类包装中恢复JSON;尚未在单一对称提取规则下完成全群组重新评分。对17个本地部署实例的存档响应审计显示,这并非其性能提升的原因:当双方使用相同提取规则时,结构化接口下的任务通过数仍为原生接口的6至10倍。仅任务通过不保证运行时或科学可行性。在此范围内,将语法生成转移至确定性软件可显著提升小型本地模型和前沿模型的可靠性。提示、评分器、响应和重新生成脚本已发布,供独立评估和扩展。

英文摘要

Scientists increasingly want to drive research software by natural-language request, but fluent output helps only if it becomes a correct machine-readable artifact. We introduce HEPToolBench, a benchmark of 28 collider-simulation tasks scored by deterministic, task-specific scorers, plus a three-task structured-debugging extension. We evaluate 42 deployments, from small locally served open-weight models to hosted frontier systems. The central experiment compares direct generation of native HEP-tool syntax with a schema-mediated interface where the model returns a typed representation that deterministic software serializes. Across five matched requests, the mean score rises from 0.418 to 0.902 and task passes from 21/210 to 159/210, with 41 of 42 deployments improving; eleven deployments, seven locally served open-weight models, go from 0/5 passes under native syntax to 5/5 under the structured interface. Values use the corrected v1.2.1 contract, in which the two native scorers that had enforced operational conventions absent from their own prompts are rescored prompt-faithfully cohort-wide. One asymmetry remains: native scorers score the recorded response verbatim and reject Markdown-fenced replies, whereas structured scorers recover JSON from such wrappers; no full-cohort rescoring under a single symmetric extraction rule has been done. An audit of archived responses from 17 locally served deployments shows this does not explain their gains: passes remain six to ten times more frequent under the structured interface when both sides use the same extraction rule. A task pass alone does not guarantee runtime or scientific viability. Within this scope, moving syntax generation into deterministic software can substantially improve reliability for both small local and frontier models. Prompts, scorers, responses, and regeneration scripts are released for independent evaluation and extension.

Comments28 pages, 8 tables, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑