发表机构
Fujitsu Limited; Waseda University; National Institute of Advanced Industrial Science and Technology(富士通株式会社; 早稻田大学; 国立研究开发法人产业技术综合研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对模仿学习对视觉模态的过度依赖,提出目标模态丢弃(TMD)方法,通过注意力估计依赖度并丢弃主导模态,结合熵正则化,在真实双臂机器人上验证了视觉缺失下的鲁棒性。
AI 中文摘要
整合多种感官模态的模仿学习策略在训练过程中容易过度依赖某一主导模态(如视觉),当该模态在推理时缺失时,可能会扰乱策略的执行。本文提出目标模态丢弃(TMD),该方法利用注意力估计对每种模态的依赖程度,并选择性地丢弃最主导的模态。同时,该方法结合了对依赖分布进行熵正则化。通过使用双臂操作器的真实机器人评估,我们表明在视觉缺失的情况下,基线策略的成功率大幅下降,而TMD能够维持任务执行。相比之下,随机选择丢弃模态且不进行熵正则化的传统丢弃方法,即使在视觉不缺失的情况下,在许多任务上也会失败。
英文摘要
Imitation learning policies that integrate multiple sensory modalities are prone to overreliance on a dominant modality, such as vision, during training, which can disrupt policy execution when that modality is lost at inference time. In this paper, we introduce Targeted Modality Dropout (TMD), in which the dependence on each modality is estimated using attention and the most dominant modality is selectively dropped. This is combined with entropy regularization over the dependence distribution. Through real-robot evaluation using a bimanual manipulator, we show that under vision loss the success rate of the baseline policy drops substantially, whereas TMD sustains task execution. In contrast, a conventional dropout that selects the dropped modality at random, without the entropy regularization, fails on many tasks even without vision loss.
Comments7 pages, 5 figures, 2 tables. Submitted to the 2027 IEEE/SICE International Symposium on System Integration (SII 2027)