arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30715cs.RO

RoboMonitor:通过预测性表示学习实现机器人任务执行的标签高效运行时监控

RoboMonitor: Label-Efficient Runtime Monitoring of Robot Task Execution via Predictive Representation Learning

Abhiroop Ajith, Gokul Narayanan, Kyle Coelho, Tingji Zhao, Yash Shahapurkar, Brian Zhu, Melih Erdogan, Ted Krubasik, Constantinos Chamzas, Eugen Solowjow

首次发表
浏览论文内容

中文总结 AI 辅助

RoboMonitor利用预测性表示学习预训练和时间监督微调,以少量标注实现高效的机器人执行监控,在阶段识别和故障检测上超越现有模型,并支持闭环部署。

中文摘要 AI 辅助

学习型机器人策略会产生动作,但仅凭其输出并不能确定执行是否按预期进行。机器人执行监控需要识别当前执行阶段、检测故障,并从执行过程中可获得的观测中识别任务完成情况。训练此类监控器所需的标注,在用于机器人策略学习的数据集中十分稀缺。我们提出了RoboMonitor,一种标签高效的视觉-语言执行监控器,它在引入监控监督之前先从这些数据集中学习。我们在涵盖12个操作任务和两种机器人形态的25小时多摄像头轨迹上进行了预训练,使用了动作条件下的未来特征预测、逆动力学和掩蔽当前预测。然后,我们将学习到的视觉和上下文编码器迁移到因果监控器,并应用时间监督微调(Temporal SFT),该方法将每个观测窗口内的监督与窗口内及重叠窗口间的一致性目标相结合。在部署时,RoboMonitor仅需任务指令和摄像头观测。在一个四项任务监控基准上,使用52个标注片段训练的RoboMonitor,在两个微调种子上实现了93.1%的平均阶段准确率和85.9%的宏召回率,超过了使用相同监控监督训练的Qwen3-VL和Robometer。其阶段准确率也超过了使用100个片段训练的Qwen3-VL和Robometer。Qwen3-VL消融实验表明,Temporal SFT将平均虚假阶段切换从15.23%降至4.95%。在闭环部署中,集成系统在40次模拟Toolbox Sorting试验中完成了39次,在40次真实Reel Packing试验中完成了35次,且未观察到虚假恢复触发。

英文摘要

Learned robot policies produce actions, but their outputs alone do not establish whether execution is progressing as intended. Robot execution monitoring requires identifying the current execution phase, detecting failures, and recognizing task completion from observations available during execution. Training such monitors requires annotations that are scarce in datasets collected for robot-policy learning. We present RoboMonitor, a label-efficient vision--language execution monitor that learns from these datasets before introducing monitoring supervision. We pre-train on 25 hours of multi-camera trajectories spanning 12 manipulation tasks and two robot embodiments, using action-conditioned future-feature prediction, inverse dynamics, and masked-present prediction. We then transfer the learned visual and context encoders to a causal monitor and apply temporal supervised fine-tuning (Temporal SFT), which combines supervision throughout each observation window with consistency objectives within and across overlapping windows. At deployment, RoboMonitor requires only the task instruction and camera observations. On a four-task monitoring benchmark, RoboMonitor trained with 52 labeled episodes achieves 93.1% mean phase accuracy and 85.9% macro recall over two fine-tuning seeds, exceeding Qwen3-VL and Robometer trained with the same monitoring supervision. Its phase accuracy also exceeds that of both Qwen3-VL and Robometer trained with 100 episodes. A Qwen3-VL ablation shows that Temporal SFT reduces mean spurious phase switching from 15.23% to 4.95%. In closed-loop deployment, the integrated system completes 39 of 40 simulated Toolbox Sorting trials and 35 of 40 real-world Reel Packing trials, with no false recovery triggers observed.

发表机构

  • Siemens Corporation(西门子股份公司)
  • ELPIS Lab, Worcester Polytechnic Institute(伍斯特理工学院ELPIS实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑