PLCWorld:闭环工厂仿真中LLM生成的PLC程序的基准测试
PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation
浏览论文内容
中文总结 AI 辅助
PLCWorld是一个闭环仿真基准,用于评估LLM生成的PLC程序,包含100个任务和473个条件对,分别报告任务成功与安全违规,发现GPT-5.5在简单任务上成功率达82.70%,困难任务仅25.10%。
中文摘要 AI 辅助
可编程逻辑控制器(PLC)通过读取传感器输入并发出控制命令来协调工业设备。评估大语言模型(LLM)生成的PLC程序是否满足任务要求和安全约束,需要观察其命令如何影响设备和工件状态。我们提出了PLCWorld,一个通用的闭环执行环境和基准测试,它将结构化文本(ST)执行与模拟的工厂响应和传感器反馈相结合。基于工业PLC程序和工程文档中识别出的控制关系,PLCWorld包含100个合成任务和473个注册的任务-条件对,涵盖运动控制和物料处理,难度由控制依赖范围定义。一个通用协议分别报告任务成功率和安全违规。验证结合了从业者审查、参考程序和替代程序、针对性反例、规范-评估器对齐检查以及与独立ST运行时的比较。参考程序和替代程序满足其适用案例,而所有542个针对性反例在至少一个注册条件下激活其指定的评估器规则。执行差距将提交档案的接受度与后续任务失败或观察到的安全违规联系起来。在所构建的任务组中,直接使用GPT-5.5在简单案例上实现了82.70%的任务成功率,但在困难案例上仅为25.10%。对六个LLM和四种改编的生成-验证工作流的评估进一步揭示了完成度、安全性和生成成本之间的差异。我们的代码、仿真环境、基准任务和基线实现可在以下URL公开获取:https://this https URL。
英文摘要
Programmable logic controllers (PLCs) coordinate industrial equipment by reading sensor inputs and issuing control commands. Evaluating whether large language model (LLM)-generated PLC programs satisfy task requirements and safety constraints requires observing how their commands affect device and workpiece states. We introduce PLCWorld, a common closed-loop execution environment and benchmark that couples Structured Text (ST) execution with simulated plant responses and sensor feedback. Grounded in control relations identified in industrial PLC programs and engineering documentation, PLCWorld contains 100 synthetic tasks and 473 registered task-condition pairs across Motion Control and Material Handling, with difficulty defined by control-dependency scope. A common protocol reports Task Success and Safety Violation separately. Validation combines practitioner review, reference and alternative programs, targeted counterexamples, specification-evaluator alignment checks, and comparisons with independent ST runtimes. Reference and alternative programs satisfy their applicable cases, while all 542 targeted counterexamples activate their designated evaluator rules under at least one registered condition. Execution Gap relates submission-profile acceptance to subsequent task failure or observed Safety Violation. Across the constructed task groups, direct GPT-5.5 achieves 82.70% Task Success on Easy cases but 25.10% on Hard cases. Evaluations of six LLMs and four adapted generation-and-verification workflows further expose differences between completion, safety, and generation cost. Our code, simulation environment, benchmark tasks, and baseline implementations are publicly available at https://yunji0516.github.io/PLCWorld/.
发表机构
- Dongguk University(东国大学)
- MOAI technologies(MOAI 科技公司)
机构由 AI 辅助整理,请以论文原文为准。