arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

凝视提示:用于视觉-语言-动作微调的时间密集人类注意力

Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning

Yihan Zhou, Rui Yan, Mingcong Li, Zheyuan Huang, Xu Yang, Xueyang Guo, Yilin Mo

arXiv 2609.34550首次发表:更新:

发表机构

Tsinghua University; Beijing Key Laboratory of Embodied Intelligence Systems; LingYu Robotics; Institute for Embodied Intelligence and Robotics, Tsinghua University(清华大学; 北京具身智能系统重点实验室; 灵宇机器人; 清华大学具身智能与机器人研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出眼动仪监督的凝视提示方法,利用VR遥操作中的凝视数据为VLA微调提供帧级视觉引导,在六个真实双臂任务上将平均成功率从26.3%提升至56.0%,并发布了含1200条轨迹的GazeMani数据集。

AI 中文摘要

视觉-语言-动作(VLA)微调在每一步都将图像与动作配对,但通常仅提供任务级别的语言指令,使得逐时刻的视觉相关性隐含不清。我们引入了眼动仪监督的凝视提示(eye-tracker-supervised gaze prompting),该方法利用VR遥操作期间记录的凝视,为VLA微调提供帧级视觉引导。在训练期间,记录的凝视位置被渲染为机器人头部相机图像上的十字准线。在部署时,一个轻量级预测器根据最近的图像和指令估计凝视位置,无需眼动仪或改变策略架构即可提供相同类型的视觉提示。以π0实例化后,凝视提示在六个真实世界双臂操作任务中将平均成功率从26.3%提高到56.0%,并且在单一策略训练于所有六个任务时也观察到了增益。我们发布了GazeMani数据集,包含1,200条带同步凝视的遥操作轨迹。

英文摘要

Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with $π_0$, gaze prompting increases mean success from $26.3\%$ to $56.0\%$ across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of $1{,}200$ teleoperated trajectories with synchronized gaze.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑