发表机构
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人操作中的故障检测,提出RAFAIL框架,通过检测实体间任务相关关系的异常来识别故障,无需故障数据,在三个真实任务上达到73.4%的平衡准确率。
AI 中文摘要
在执行过程中检测故障对于可靠的机器人操作至关重要。视觉语言模型(VLM)能够从语义上评估任务结果,但会增加运行时计算开销,而分布外(OOD)检测器可能对无害的场景变化做出响应,而非与故障相关的偏差。我们提出了RAFAIL,一个用于检测机器人操作过程中执行故障的框架。RAFAIL通过检测实体之间任务相关关系中的异常来识别故障,例如夹爪与物体之间或物体与其目标之间的关系。通过将OOD检测聚焦于观察中相关的部分,RAFAIL降低了对任务无关场景变化的敏感性。在离线阶段,VLM为成功演示标注任务进度和关系重要性,用于学习基于点云的关系表示,而不依赖策略内部特征。在运行时,特定于关系的OOD检测器评估这些表示,同时在不进行VLM推理的情况下预测关系重要性和任务进度。RAFAIL无需故障数据,在三个真实世界机器人操作任务中实现了73.4%的平衡准确率,优于所评估的最强OOD和基于不确定性的基线方法。
英文摘要
Detecting failures during execution is essential for reliable robotic manipulation. Vision-language models (VLMs) can assess task outcomes semantically but add runtime computation, whereas out-of-distribution (OOD) detectors may respond to harmless scene variations rather than failure-relevant deviations. We introduce RAFAIL, a framework for detecting execution failures during robotic manipulation. RAFAIL identifies failures by detecting anomalies in task-relevant relationships between entities, such as a gripper and an object or an object and its target. By focusing OOD detection on relevant parts of the observation, RAFAIL reduces sensitivity to task-irrelevant scene variation. Offline, a VLM annotates successful demonstrations with task progress and relationship importance, which are used to learn point-cloud-based relationship representations without relying on policy-internal features. At runtime, relationship-specific OOD detectors evaluate these representations while relationship importance and task progress are predicted without VLM inference. RAFAIL requires no failure data and achieves 73.4% balanced accuracy across three real-world robotic manipulation tasks, outperforming the strongest evaluated OOD- and uncertainty-based baselines.