发表机构
Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST)(日本产业技术综合研究所人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对端到端视觉运动策略中人类难以理解和校正视觉注意力的问题,提出引导注意力框架,通过预测任务相关注意力关键点并结合扩散动作策略,经用户校正后由跟踪模块传播,提升了机器人在OOD条件下的操纵性能。
AI 中文摘要
端到端视觉运动策略让人类几乎没有机会理解或校正策略的视觉注意力。我们提出了引导注意力,这是一个视觉运动模仿学习框架,引入可解释和可校正的视觉注意力作为显式中间表示。从相机图像预测与任务相关的注意力关键点,并调节基于扩散的动作策略。用户可在部署初始化时检查并可选地校正选定关键点,之后校正后的注意力由跟踪模块在整个执行过程中自动传播。仿真和现实世界实验表明,引导注意力能持续提高机器人操纵性能,尤其在位置和外观分布外(OOD)条件下。
英文摘要
End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a tracking module. Experiments in simulation and the real world demonstrate that GuidedAttention consistently improves robot manipulation performance, particularly under positional and appearance out-of-distribution (OOD) conditions. https://mmurooka.github.io/guided-attention-project-page
CommentsProject page added