发表机构
Research & Development Group, Hitachi, Ltd.(日立制作所研究开发集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FRAM通过预测轨迹引导视觉特征选择,以138.7M参数在LIBERO上达到92.2%成功率,接近3.3B参数的π0,验证了未来运动信息对小型策略的有效性。
AI 中文摘要
视觉-语言-动作模型在机器人操作中表现出强大的性能,但通常需要大量的参数。在这项工作中,我们提出了未来表示动作模型(FRAM),这是一个小型策略,明确地将未来的末端执行器轨迹与当前的视觉输入联系起来。FRAM使用预测轨迹的图像坐标作为空间指针,并从当前图像中读取与运动相关的局部视觉特征。这将用于动作生成的信息组织为参考位置(Where)、视觉状态(What)和未来运动(Future)。轨迹标签从演示和相机几何中自动生成,因此无需手动标注。凭借138.7M参数(包括冻结的语言编码器),FRAM在四个标准LIBERO套件上达到了92.2%的平均成功率,接近具有3.3B参数的π0的94.2%。无需额外训练,它还在LIBERO-Plus上达到了67.3%的平均成功率。消融实验证实,未来轨迹和局部视觉特征均能提高性能和鲁棒性。在真实的双臂UR5e上,FRAM仅使用腕部相机即可叠放杯子,包括选择并在左右臂之间切换。这些结果表明,基于未来运动选择视觉信息是在小型机器人策略中获得高性能和鲁棒性的有效方法。
英文摘要
Vision-Language-Action models achieve strong performance in robot manipulation, but often require large numbers of parameters. In this work, we propose the Future Representation Action Model (FRAM), a small policy that explicitly links the future end-effector trajectory to the current visual input. FRAM uses the image coordinates of the predicted trajectory as spatial pointers and reads local visual features related to the motion from the current image. This organizes the information for action generation into the reference position (Where), the visual state (What), and the future motion (Future). Trajectory labels are generated automatically from demonstrations and camera geometry, so no manual annotation is needed. With 138.7M parameters, including a frozen language encoder, FRAM reaches an average success rate of 92.2% over the four standard LIBERO suites, close to the 94.2% of $π_0$ with 3.3B parameters. Without extra training, it also reaches an average of 67.3% on LIBERO-Plus. Ablations confirm that both the future trajectory and the local visual features improve performance and robustness. On a real dual-arm UR5e, FRAM stacks cups using only wrist cameras, including choosing and switching between the left and right arms. These results show that selecting visual information based on future motion is an effective way to obtain both high performance and robustness in a small robot policy.