规划还是学习:多资产维护中的可靠性与成本
Planning or Learning: Reliability and Cost in Multi-Asset Maintenance
- Industrial AI Lab, Hitachi America Ltd.(日立美国有限公司工业人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究实证比较规划与强化学习在多资产维护中的表现,发现规划以硬约束实现零故障,RL在低惩罚下成本更低但存在非零故障,并探讨了约束机制,表明两者互补。
AI中文摘要:
工业维护系统涉及多个相互作用的资产和共享资源,这使得使用单一决策框架来平衡可靠性和运营成本颇具挑战。尽管近期工作聚焦于将强化学习(RL)用于维护调度,但在相同设置下与规划方法的直接比较仍然有限。在本研究中,我们利用运行至失效数据,对多资产轴承维护中的规划与RL方法进行了实证比较。我们考察了在一系列故障惩罚场景下,这些方法在平衡预防性维护与可容忍故障时的行为表现。我们观察到由目标函数设定驱动的一致行为差异。规划将可靠性作为硬约束,产生零故障策略,其总成本对故障惩罚的幅度基本不敏感。RL智能体优化期望成本,并常随惩罚变化在预防性维护与偶发故障之间进行权衡,导致在低惩罚制度下成本较低,但即使惩罚较高时仍持续存在非零故障。我们还研究了轻量级约束机制,包括奖励塑形和动作掩蔽,以鼓励RL的可靠性。从实践角度看,当需要严格可靠性且部署周期较短时,规划可能更合适;而当可接受有限故障且优先考虑长期运营效率时,RL可能提供成本效益高的策略。总体而言,本研究阐明了多资产维护中可靠性与成本之间的权衡,并表明规划与RL是互补的方法。除这些发现外,受控基准协议本身统一了跨范式的环境、成本模型和评估,为在其他维护设置中比较决策方法提供了可复用的模板。
英文摘要:
Industrial maintenance systems involve multiple interacting assets and shared resources, making it challenging to balance reliability and operational cost using a single decision framework. While recent work has focused on reinforcement learning (RL) for maintenance scheduling, direct comparisons with planning approaches under identical settings remain limited. In this work, we empirically compare planning and RL for multi-asset bearing maintenance using run-to-failure data. We examine how these methods behave when balancing preventive maintenance against tolerable failures across a range of failure penalty scenarios. We observed a consistent behavioral difference driven by objective formulation. Planning enforces reliability as a hard constraint and produces zero-failure policies whose total cost is largely insensitive to the magnitude of failure penalties. RL agents optimize expected cost and often trade off preventive maintenance against occasional failures as penalties vary, resulting in lower costs under low-penalty regimes but persistent non-zero failures even when penalties are high. We also investigate lightweight constraint mechanisms, including reward shaping and action masking, to encourage RL's reliability. From a practical perspective, planning may be more suitable when strict reliability is required and deployment horizons are short, whereas RL may provide cost-efficient policies when limited failures are acceptable and long-run operational efficiency is prioritized. Overall, this study clarifies the trade-offs between reliability and cost in multi-asset maintenance and suggests that planning and RL are complementary approaches. Beyond these findings, the controlled benchmark protocol itself that unifies environment, cost model, and evaluation across paradigms, offers a reusable template for comparing decision-making approaches in other maintenance settings.