AI 中文总结
该研究针对IM技术依赖全序轨迹的局限,提出将其从全序轨迹提升至偏序轨迹的方法,避免线性化爆炸,提升了过程发现的效率与准确性。
AI 中文摘要
归纳式矿工(Inductive Miner,IM)家族是一类杰出的过程发现技术,将高效的递归分解与构造性的合理性保证相结合。然而,IM技术通常假定轨迹是活动发生的全序序列。这一假定虽方便,却可能引入系统性偏差:活动可能具有持续时间、事件可能共享粗略时间戳,或数据仅约束部分事件对。将此类执行强行转换为任意序列会隐藏固有并发性,还可能引入从未作为因果约束被观察到的顺序依赖关系。偏序提供了更忠实的表示,但将其整合到IM发现中颇具挑战性,因为标准抽象是基于序列的;直接复用它们需要对每个偏序进行线性化,而在高并发性下这会变得极其昂贵。我们提出将IM发现从全序轨迹提升至偏序轨迹。我们未重新设计矿工及其割检测逻辑,而是重新定义轨迹抽象层与递归投影,使其直接作用于偏序。该方法对全序轨迹具有保守性,避免了线性化爆炸,保留了IM具有吸引力的递归结构与保证。实验结果表明,所提出的提升方法避免了线性化的组合开销,降低了对带时间戳事件数据中任意平局打破的敏感性,且通过在轨迹级别保留并发性,可从更少的观察中学习过程行为。
英文摘要
The Inductive Miner (IM) family is a prominent class of process discovery techniques, combining efficient recursive decomposition with soundness-by-construction guarantees. However, IM techniques usually assume traces to be totally ordered sequences of activity occurrences. This assumption is convenient, but can introduce systematic bias: activities may have durations, events may share coarse timestamps, or the data may constrain only some event pairs. Forcing such executions into arbitrary sequences hides inherent concurrency and may introduce sequential dependencies that were never observed as causal constraints. Partial orders provide a more faithful representation, but integrating them into IM discovery is challenging because standard abstractions are sequence-based; directly reusing them would require linearizing each partial order, which becomes prohibitively expensive under high concurrency. We introduce a lifting of IM discovery from total orders to partially ordered traces. Instead of redesigning the miner and its cut detection logic, we redefine the trace abstraction layer and the recursive projections to operate directly on partial orders. The approach is conservative over totally ordered traces, avoids linearization explosion, and preserves the recursive structure and guarantees that make IM attractive. Experimental results show that the proposed lifting avoids the combinatorial overhead of linearization, reduces sensitivity to arbitrary tie-breaking in timestamped event data, and allows process behavior to be learned from fewer observations by preserving concurrency at the trace level.