PHITSBench:用于使用自然语言生成人工智能辅助的PHITS辐射传输输入的执行评分基准
PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language
浏览论文内容
中文总结 AI 辅助
介绍用于PHITS辐射传输输入生成的PHITSBench基准,涵盖三类任务。评估五种基于GPT - 5.4的配置,无特定知识时模型在编辑和修复任务表现好,Reproduce任务差。知识目录和智能体执行可提升成功率,结果显示建模进展依赖多方面因素。
中文摘要 AI 辅助
我们引入了PHITSBench,这是一个用于蒙特卡罗粒子与重离子输运代码系统(PHITS)的执行评分基准。PHITSBench包含282个可进行传输评分的任务,涵盖三个常见工作流程类别:参数编辑(Edit)、语法修复(Repair)以及根据自然语言描述生成完整模拟(Reproduce)。每个任务使用综合指标分数进行评估,该分数结合了执行成功率以及生成的与参考传输可观测量之间的一致性。我们使用PHITSBench评估了五种基于GPT - 5.4的配置,从零样本提示到知识增强和智能体工作流程。在没有特定领域知识的情况下,模型在编辑和修复任务上表现良好(分别为95%和70%的成功率),但无法从头生成正确的模拟(在Reproduce任务上成功率为0%)。随用户手册提供的结构化、机器可读的PHITS知识目录将单次Reproduce任务的成功率提高到了57%。智能体执行进一步提高到66 - 73%,但计算成本增加。故障分析表明,剩余错误主要由物理可观测量的错误选择和配置而非语法生成主导。这些结果表明,人工智能辅助辐射传输建模的未来进展将同样依赖于机器可读知识库、精心策划的领域训练数据集和基于执行的评估环境,以及基础模型本身的进展。
英文摘要
We introduce PHITSBench, an execution-scored benchmark for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS). PHITSBench comprises 282 transport-scorable tasks spanning three common workflow categories: parameter editing (Edit), syntax repair (Repair ), and complete simulation generation from natural-language descriptions (Reproduce). Each task is evaluated using a Composite Metric Score that combines execution success with agreement between generated and reference transport observables. Using PHITSBench, we evaluate five GPT-5.4-based configurations ranging from zero-shot prompting to knowledge-augmented and agentic workflows. Without domain-specific knowledge, the model performs well on editing and repair tasks (95% and 70% success, respectively) but fails to generate correct simulations from scratch (0% success on the Reproduce track). A structured, machine-readable PHITS knowledge catalog, supplied alongside the user manual, raises single-shot Reproduce-task success to 57%. Agentic execution provides a further improvement to 66-73%, but at increased computational cost. Failure analysis shows that the remaining errors are dominated by incorrect selection and configuration of physical observables rather than syntax generation. These results suggest that future progress in AI-assisted radiation-transport modeling will depend as much on machine-readable knowledge bases, curated domain-training datasets, and execution-grounded evaluation environments as on advances in foundation models themselves.
发表机构
- MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。