多模态机器人演示中指令-轨迹不匹配的审计
Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations
- AI Robot Association (AIRoA)(人工智能机器人协会(AIRoA))
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态机器人演示中指令-轨迹不匹配问题,提出无需训练的MMPF审计框架,在LIBERO基准及真实机器人数据上实现最优ITM检测与标签修正,可提升下游策略学习性能并展示过滤演示的权衡。
AI中文摘要:
用于训练视觉-语言-动作策略的机器人演示数据集可能包含一种微妙但有害的失效模式:行为正确的轨迹却与错误的语言指令配对。我们研究对这些指令-轨迹不匹配(Instruction-Trajectory Mismatches,ITMs)的事后审计。与失败的 rollout 不同,ITMs 通常看起来合理,会破坏策略学习到的语言-行为映射。我们提出多模态概率融合(Multimodal Probabilistic Fusion,MMPF),这是一种无需训练的审计框架,将每个模态视为一个专家,从局部邻域一致性和全局原型相似性估计任务标签分布,再通过乘积专家框架中基于预测熵的权重融合各模态。在注入指令不匹配的 LIBERO 基准测试及带噪声的真实机器人数据上,MMPF 实现了最强的整体 ITM 检测和标签修正准确率。我们还表明,在需要语言来区分任务的场景中,审计能提升多数下游策略学习效果。真实机器人实验显示,与重新标注相比,我们的方法可提升策略性能,并展示了过滤演示的权衡关系。
英文摘要:
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. We study post-hoc auditing of these Instruction-Trajectory Mismatches (ITMs). Unlike failed rollouts, ITMs often look plausible, and can corrupt the language-behavior mapping learned by the policy. We propose Multimodal Probabilistic Fusion (MMPF), a training-free auditing framework that treats each modality as an expert, estimates a task-label distribution from local neighborhood agreement and global prototype similarity, and then fuses modalities with predictive-entropy weighting in a product of experts. Across LIBERO benchmarks with injected instruction mismatches and noisy real-robot data, MMPF achieves the strongest overall ITM detection and label correction accuracy. We also show that auditing improves most downstream policy learning in settings where language is needed to disambiguate the task. We demonstrate in real robot experiments that our method can achieve improved policy performance and show the trade-off of filtering demonstrations compared to relabeling.