发表机构
The University of Auckland; Fudan University(奥克兰大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出帧级时间对齐(FLTA)框架,利用全局进展和局部时间顺序两种先验,无需帧级标注即可对齐人机演示,在仿真和真实操作任务中显著提升成功率。
AI 中文摘要
将基于人类视频预训练的视觉表征迁移到机器人操作任务中,需要学习人类演示与机器人演示之间可靠的对应关系。然而,成对演示在执行速率以及不直接反映任务进展的非关键帧比例上可能存在差异。因此,相同相对时间戳的帧可能代表不同的任务阶段,这可能导致对应关系学习失败。为解决这些问题,我们提出了帧级时间对齐(FLTA),这是一个利用两种时间先验来适配预训练于人类视频的视觉编码器以用于机器人操作的框架。它无需帧级对应标注即可学习共享的任务进展表征,允许不同相对时间位置的帧进行匹配。全局进展先验将归一化时间位置与视觉相似性相结合,构建软对应目标。局部时间顺序先验惩罚向后转换,同时允许停留和变化的前向速率以适应执行速率的差异。使用ResNet-50和ViT编码器,我们的方法在平均仿真成功率上相较于最佳基线分别实现了46.93%和65.96%的相对提升,并在真实世界操作任务中实现了更高的任务成功率。这些结果还表明,有效的人机适配更多地取决于更新哪些参数,而非更新参数的数量。我们的项目页面可从此https URL访问。
英文摘要
Transferring visual representations pretrained on human videos to robot manipulation requires learning reliable correspondences between human and robot demonstrations. However, paired demonstrations can differ in execution rate and in the proportion of non-key frames that do not directly reflect task progress. Frames at the same relative timestamp may therefore represent different task stages, which can cause correspondence learning to fail. To address these issues, we propose Frame-Level Temporal Alignment (FLTA), a framework that uses two temporal priors to adapt visual encoders pretrained on human videos for robot manipulation. It learns shared task-progress representations without frame-level correspondence annotations, allowing frames at different relative temporal positions to match. A global progress prior combines normalized temporal positions with visual similarity to construct a soft correspondence target. A local temporal order prior penalizes backward transitions while allowing stays and varying forward rates to accommodate execution-rate differences. With ResNet-50 and ViT encoders, our method achieves relative improvements of 46.93% and 65.96%, respectively, over the best baselines in average simulation success rates and achieves higher task success rates on real-world manipulation tasks. These results also suggest that effective human-robot adaptation depends less on the number of parameters updated than on which parameters are selected. Our project page is available at https://rtx5090ultra.github.io/FLTA-Project-Page/.
Comments24 pages, 10 figures, 12 tables