发表机构
Samsung Robotics eXperience; National University of Singapore(三星机器人体验中心; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLA模型在高精度操作中注意力分散的问题,提出ActGaze方法,利用反事实视觉干预从动作目标中学习注视监督,实验证明其提升注意力聚焦并超越基线。
AI 中文摘要
当前的视觉-语言-动作(VLA)模型通常难以应对高精度机器人操作任务。我们将这一局限主要归因于其视觉注意力分散在与任务无关的区域。为解决此问题,我们提出ActGaze,一种训练方法,引导VLA策略将注视集中于任务相关区域,类似于人类在执行精确动作时注视关键视觉线索。与先前依赖外部标签进行注视监督的方法不同,ActGaze通过使用反事实视觉干预直接从VLA自身的动作目标中导出空间监督,以识别对动作预测至关重要的区域。在四个高精度机器人操作任务上进行的大量真实机器人实验表明,ActGaze能诱导对任务相关区域更集中的视觉注意力,并持续优于基础VLA策略及其他视觉接地方法。
英文摘要
Current Vision-Language-Action (VLA) models often struggle with high-precision robotic manipulation. We attribute this limitation primarily to their visual attention being dispersed across task-irrelevant regions. To address this issue, we propose ActGaze, a training approach that guides VLA policies to gaze on task-relevant regions, much like humans gaze on critical visual cues while executing precise movements. Unlike prior methods that rely on external labels for gaze supervision, ActGaze derives spatial supervision directly from the VLA's own action objective by using counterfactual visual interventions to identify regions that are critical for action prediction. Extensive real-robot experiments on four high-precision robotic manipulation tasks demonstrate that ActGaze induces more focused visual attention on task-relevant regions and consistently outperforms the base VLA policy and other visual-grounding approaches.
Comments11 pages, 7 figures