发表机构
Amazon FAR; UC Berkeley; Stanford University(亚马逊FAR; 加州大学伯克利分校; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EyeRobot 2.0提出主动注视框架,仅用单立体相机实现精细双手操控,通过分层强化学习训练注视伺服与目标选择,在真实和模拟任务中显著超越被动立体视觉,并匹配或超越腕部相机策略。
AI 中文摘要
受人类视觉启发,我们提出了一种利用主动注视实现仅凭单个立体相机即可完成精细双手操控的框架。EyeRobot 2.0通过旋转两个眼球视角使其注视中心对准场景中的三维注视点,从而在物理上关注该点。生成的图像通过将更多视觉令牌分配给图像中心进行中央凹处理,将计算集中于任务相关特征。这种主动视觉固定(AVF)要求在任务执行过程中进行精心协调的注视,我们通过分层方式实现:首先训练一个以目标物体为条件的低级注视伺服策略,然后训练一个根据任务进度发出注视目标的目标选择器。两个模块均在真实世界数据上使用强化学习训练:第一个使用密集几何奖励训练,第二个与BC夹爪策略协同训练,使其能够发现与人类执行任务时的注视序列相似的注视序列。EyeRobot 2.0进一步利用注视,将夹爪信息规范化为注视相对SE(3)坐标系,从而压缩了待学习的动作分布规模。我们收集了7个真实世界任务和6个模拟任务的遥操作数据,并进行了超过1000次物理机器人试验和1800次模拟机器人试验,将EyeRobot 2.0与在相同数据上训练的被动立体视觉策略及自我+腕部相机策略进行比较。移除腕部相机对标准策略代价高昂:仅使用被动立体视觉时,真实世界成功率从52%降至27%。EyeRobot 2.0仅凭立体视觉弥补了这一差距,在真实环境中比被动立体视觉高出40%,在模拟环境中高出20%。当腕部视野清晰时,它达到与自我+腕部策略相当的水平(69%对64%),而当抓取物体遮挡腕部相机时,其成功率是后者的两倍以上(48%对22%)。
英文摘要
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human's fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)
CommentsProject Page: https://eyerobot2.github.io/