发表机构
Deakin University(迪肯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机械故障检测依赖监督学习且面临故障标签稀缺的问题,提出对抗性逆强化学习框架,直接从状态转换恢复“健康”奖励,无需人工设计和故障标签,在三个基准测试中表现出色,实现非饱和检测后一致性。
AI 中文摘要
机械故障检测严重依赖监督学习,在实际场景中面临故障标签稀缺的问题。强化学习虽能为退化的顺序性建模,但当前基于强化学习的机械故障检测方法将问题简化为静态上下文博弈,忽略状态转换和时间折扣因子。本文提出对抗性逆强化学习框架,将机械故障检测视为离线逆强化学习问题。该方法直接从观测到的状态转换中恢复内在的“健康”奖励,无需人工奖励设计和故障标签。在三个运行至故障的基准测试中,该方法是唯一在所有数据集上实现非饱和检测后一致性的方法,而基于上下文博弈的基线无法检测到逐渐退化,重建模型则陷入总是异常的状态。
英文摘要
Machinery fault detection (MFD) remains heavily reliant on supervised learning, which struggles with the scarcity of fault labels in real-world settings. While reinforcement learning (RL) offers a framework to model the sequential nature of degradation, current ``RL-based'' MFD methods reduce the problem to a static contextual bandit (CB) formulation: by ignoring state transitions and discarding the temporal discount factor, they collapse to standard supervised classification. We propose an adversarial inverse reinforcement learning (AIRL) framework that treats MFD as an offline IRL problem. Unlike reconstruction-based approaches that rely on static error margins, or CBs that ignore dynamics, our method recovers an intrinsic "health" reward directly from observational state transitions, requiring neither manual reward engineering nor fault labels. On three run-to-failure benchmarks (HUMS2023, IMS, XJTU-SY), AIRL is the only method achieving non-saturated post-detection consistency across all datasets, while CB baselines fail to detect gradual degradation and reconstruction models collapse into always-anomalous states. Code and data: https://github.com/dhirajneupane/AIRL-MFD-DN.