发表机构
Tsinghua University; Beijing Key Laboratory of Embodied Intelligence Systems; LingYu Robotics; Institute for Embodied Intelligence and Robotics, Tsinghua University(清华大学; 北京具身智能系统重点实验室; 灵宇机器人; 清华大学具身智能与机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出眼动仪监督的凝视提示方法,利用VR遥操作中的凝视数据为VLA微调提供帧级视觉引导,在六个真实双臂任务上将平均成功率从26.3%提升至56.0%,并发布了含1200条轨迹的GazeMani数据集。
AI 中文摘要
视觉-语言-动作(VLA)微调在每一步都将图像与动作配对,但通常仅提供任务级别的语言指令,使得逐时刻的视觉相关性隐含不清。我们引入了眼动仪监督的凝视提示(eye-tracker-supervised gaze prompting),该方法利用VR遥操作期间记录的凝视,为VLA微调提供帧级视觉引导。在训练期间,记录的凝视位置被渲染为机器人头部相机图像上的十字准线。在部署时,一个轻量级预测器根据最近的图像和指令估计凝视位置,无需眼动仪或改变策略架构即可提供相同类型的视觉提示。以π0实例化后,凝视提示在六个真实世界双臂操作任务中将平均成功率从26.3%提高到56.0%,并且在单一策略训练于所有六个任务时也观察到了增益。我们发布了GazeMani数据集,包含1,200条带同步凝视的遥操作轨迹。
英文摘要
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with $π_0$, gaze prompting increases mean success from $26.3\%$ to $56.0\%$ across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of $1{,}200$ teleoperated trajectories with synchronized gaze.