SeeTraceAct: 跨具身演示视频中的可见性感知潜在规划
SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos
浏览论文内容
中文总结 AI 辅助
提出SeeTraceAct框架,通过可见性感知的未来末端执行器轨迹预测增强空间定位,实现基于单次跨具身演示视频的机器人策略泛化,在模拟和真实场景中取得最优成功率。
中文摘要 AI 辅助
视觉-语言-动作模型(VLA)是有前途的通用机器人策略,但将其适应新任务通常需要昂贵的任务特定遥操作数据。作为替代,我们研究一次性演示条件VLA,其中机器人策略以未见任务的单个演示视频为条件。我们发现,当成功执行需要精确定位小目标区域时,现有的端到端方法往往难以应对。为解决这一限制,我们提出SeeTraceAct,一种演示条件VLA框架,通过可见性感知的未来末端执行器轨迹预测来鼓励精确的空间定位。为实现跨具身演示的可重复评估,我们引入并发布了RoboCasa-DC,这是RoboCasa的演示条件扩展,包含成对的人形机器人视频。在RoboCasa-DC和真实世界基准(Franka Panda臂以人类演示为条件)上的实验表明,SeeTraceAct优于基线,在所有四个RoboCasa-DC设置中实现了最佳成功率,并将真实世界平均成功率提高了12.5个百分点。
英文摘要
Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-specific teleoperation data. As an alternative, we study one-shot demo-conditioned VLAs, where a robot policy is conditioned on a single demonstration video of an unseen task. We find that existing end-to-end approaches often struggle when successful execution requires precisely localizing small target regions. To address this limitation, we propose SeeTraceAct, a demo-conditioned VLA framework that encourages precise spatial grounding through visibility-aware prediction of future end-effector traces. To enable reproducible evaluation with cross-embodiment demonstrations, we introduce and release RoboCasa-DC, a demo-conditioned extension of RoboCasa with episode-paired humanoid videos. Experiments on RoboCasa-DC and a real-world benchmark, where a Franka Panda arm is conditioned on human demonstrations, show that SeeTraceAct outperforms baselines, achieving the best success rate across all four RoboCasa-DC settings and improving real-world average success by 12.5 percentage points.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- Allen Institute for AI(Allen人工智能研究所)
- Johns Hopkins University(约翰霍普金斯大学)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。