arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用智能体AI构建流程建模工具:PM4Py-UCM的经验报告

Building a Process-Modeling Tool using Agentic AI: An Experience Report on PM4Py-UCM

Daniel Amyot

arXiv 2607.28825首次发表:更新:

AI 中文总结

本文通过AI辅助构建开源流程建模工具PM4Py-UCM的深入案例,提出相关工具包与分类法,总结了开发经验教训及验证策略。

AI 中文摘要

企业建模(EM)工具通常复杂且难以扩展,而用户可能希望探索当前不存在的新EM功能和能力。AI编码智能体可通过支持开发新功能和完整工具在此提供帮助,但大型语言模型(LLM)主导编写的建模语言工具是否可信仍是一个问题。本文报告了AI辅助构建PM4Py-UCM的过程,这是一款从事件日志挖掘用例图(UCM)模型的开源工具。PM4Py-UCM的功能包括流程挖掘工具的常见功能(如性能热图和仪表板)以及独特功能(如挖掘可执行场景/变体和模型分解)。我们挖掘了开发记录,该记录包含18次智能体会话(65小时内374次人工交互和10328次工具操作)、151次提交、20次发布以及从108个测试函数增长到691个测试函数的测试套件,以通过单个深入案例描述该工具如何使用智能体(Claude Code)构建,并辅以对生成代码的独立静态评估(涵盖覆盖率、复杂性、可维护性、安全性、架构)。我们提供了可复现、隐私保护的工具包和分类法,用于对人工交互进行分类并标记跨领域一致性工作、智能体修正和撤回的请求。截至0.7.4版本,修复数量与功能数量之比为2.3:1,约18%的交互用于修正智能体错误。功能迭代带来了可测量的文档/测试/笔记本一致性工作滞后,测试与功能同步增长。我们最终提出了经验教训,核心是使模型转换可机械检查,以及基于预言机的验证策略,该策略弥合了“智能体称其有效”的差距,以负责任地使用AI工程化EM工具。

英文摘要

Enterprise-modeling (EM) tools are often complex and hard to extend. Yet, users may want to explore new EM features and capabilities that currently do not exist. AI coding agents can help here by enabling the development of new capabilities and entire tools, but whether we can trust a modeling-language tool an LLM largely wrote remains a question. This paper reports on the AI-assisted construction of PM4Py-UCM, an open-source tool that mines Use Case Map (UCM) models from event logs. PM4Py-UCM's capabilities include some expected from process mining tools (e.g., performance heat-maps and dashboards) and distinctive ones (e.g., mined executable scenarios/variants, and model decomposition). We mined the development record itself, composed of 18 agent sessions (374 human turns and 10,328 tool actions over 65 hours), 151 commits, 20 releases, and a test suite grown from 108 to 691 test functions, in order to characterize, in a single in-depth case, how the tool was built with an agent (Claude Code), complemented by an independent static assessment of the resulting code (coverage, complexity, maintainability, security, architecture). We contribute a reproducible, privacy-preserving toolkit and taxonomy that classify human turns and flag cross-cutting consistency work, agent corrections, and retracted requests. Up to version 0.7.4, fixes outnumber features 2.3:1, with ~18% of turns for correcting agent errors. Feature waves dragged a measurable tail of documentation/test/notebook consistency work, and tests grew lockstep with features. We finally present lessons learned, centered on making model transformations mechanically checkable, and the oracle-based validation strategy that closed the "the agent said it works" gap, for responsibly engineering EM tooling with AI.

Comments17 pages, 9 figures, 4 tables, submitted to a conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑