arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SemaPLC:一种基于项目、经验证门控的PLC代码生成智能体框架

SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation

Yanlun Tu, Huacan Wang, Ziyue Zhou, Jie Zhou, Ningyan Zhu, Ge Chen, Wangyi Chen, Tengfei Zhou, Yifan Zhou, Dasheng Yang, Xiaofeng Mou, Hui Zhang, Yi Xu

arXiv 2608.18565首次发表:更新:

AI 中文总结

SemaPLC是基于项目、经验证门控的PLC代码生成智能体框架,通过外部检查确认任务完成,在多模型的POU任务及项目上下文任务中均表现最优,动态行为测试得分显著高于基线。

AI 中文摘要

可编程逻辑控制器(PLC)用于运行工业工厂,大型语言模型现已能为其生成独立的程序组织单元(POU)。此类逻辑是否能集成到现有PLC项目并正确运行,此前仅通过有限测试进行过检查。我们提出SemaPLC,这是一种基于项目、经验证门控的智能体框架,由常规工具组装而成,但受严格完成规则管控。SemaPLC不会在模型判定自身输出足够时停止,仅当记录的外部检查确认任务完成时,才会宣布任务完成。这些检查涵盖规范、编译以及在实时运行时的行为。在与现有基准匹配的117个独立POU任务中,它在所有7个模型上均达到最高的严格验证通过率(均值为72.6%)。在65个生成逻辑必须在真实项目内编译并运行的项目上下文任务中,它在集成编译、静态行为和动态行为方面均达到最高均值。在三个层级中,动态行为最具揭示性,我们通过将生成的逻辑和参考逻辑部署到实时PLC运行时并比较其执行轨迹来对其进行测量。所有方法的静态得分均相差在10分以内,而动态得分则将它们明显区分:基线的动态得分介于22.4至31.4之间,而SemaPLC的动态得分达52.2。总体而言,我们的验证门控框架提升了每个层级的均值,在运行时的提升最为显著。执行而非静态评分是判断生成的控制逻辑是否真正有效的可靠测试。SemaPLC已开源,网址为this https URL。

英文摘要

Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independent program organization units (POUs) for them. Whether such logic integrates into an existing PLC project and then runs correctly has been checked only in limited tests. We present \textsc{SemaPLC}, a project-grounded and verification-gated agent harness assembled from conventional tools but governed by a strict completion rule. Rather than stopping when the model judges its own output adequate, \textsc{SemaPLC} declares a task complete only when logged external checks confirm it. Those checks cover the specification, the compilation, and the behavior on a live runtime. On 117 independent-POU tasks matching existing benchmarks, it attains the highest strict verified pass rate on all seven models (72.6\% mean). On a project-context track of 65 tasks whose generated logic must compile and run inside a real project, it attains the highest mean on integrated compilation, static behavior, and dynamic behavior. Of the three layers, dynamic behavior is the most revealing. We measure it by deploying the generated and the reference logic to a live PLC runtime and comparing their executed traces. All methods fall within 10 static points of one another, whereas dynamic scores separate them sharply, from 22.4 to 31.4 for the baselines against 52.2 for \textsc{SemaPLC}. Overall, our verification-gated harness raises the mean at every layer and most sharply at runtime. Execution, not static scoring, is the faithful test of whether generated control logic actually works. \textsc{SemaPLC} is open-sourced at https://github.com/midea-ai/SemaPLC.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑