什么情况下关键?诊断和改进视觉运动模仿策略中的条件视觉接地
What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies
浏览论文内容
中文总结 AI 辅助
该研究针对视觉运动模仿策略在引入相似物体时的目标选择失效问题,以ACT为基础提出三类互补干预措施,在仿真和物理UR3e机器人上提升了鲁棒性,还验证了其在预训练视觉-语言-动作策略上的通用性。
中文摘要 AI 辅助
视觉运动模仿策略在分布内视觉条件下可实现高性能,但引入视觉相似的物体或容器时会失效。我们将此行为视为条件视觉接地问题:成功控制所需的视觉目标会随操作阶段变化,在更复杂的任务中还会随观测到的任务状态变化。使用基于Transformer的动作分块(ACT),我们系统引入具有可控颜色和形状相似度的干扰物体和容器,并将失效定位到抓取和放置环节。我们发现干扰敏感度对视觉相似度类型和操作阶段均具有特异性。基于此诊断,我们评估干扰增强、阶段相关注意力正则化和基于外观的视觉提示作为互补干预措施,以在保留控制所需空间信息的同时改进目标选择。这些干预措施在仿真环境和物理UR3e机器人上均显著提升了鲁棒性。我们还在一个预训练的视觉-语言-动作策略上检查了相同的失效模式,该策略用于状态条件下的器械操作任务,其中医疗器械的观测状态决定了正确目标位置。总体而言,结果表明,即使底层操作技能完好,视觉干扰也会导致错误的物体或目标位置选择,而显式改进目标选择可在不同视觉运动策略学习机制中大幅恢复性能。
英文摘要
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.
发表机构
- Fraunhofer Institute for Production Systems and Design Technology IPK(弗劳恩霍夫生产系统与设计技术研究所IPK)
- Technische Universität Berlin(柏林工业大学)
机构由 AI 辅助整理,请以论文原文为准。