发表机构
College of Computer Science and Artificial Intelligence, Fudan University; Singapore Management University; Institute of Trustworthy Embodied AI, Fudan University(复旦大学计算机科学与人工智能学院; 新加坡管理大学; 复旦大学可信具身智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SpatialHarness是一种无需微调策略或更改传感设置的测试时具身支架,通过同步模拟场景提升空间可观测性,使GPT-6 Astra在四项真实机器人精细操作任务上的成功率显著提升。
AI 中文摘要
前沿多模态基础模型(如GPT-6 Astra)近期展现出直接控制机器人的强大潜力,但其在精细操作任务上的性能仍有限。我们认为,性能不佳的重要原因未必是策略能力不足,而是空间可观测性不足——现有物理相机设置可能无法充分呈现任务关键的空间关系。我们提出SpatialHarness,这是一种测试时具身支架,无需对策略进行微调或更改物理传感设置,即可为精细机器人操作提供测试时空间支架。SpatialHarness会维护与真实世界执行同步的在线模拟场景,识别任务关键的空间关系,并渲染互补的虚拟视图,将这些关系呈现给冻结的多模态策略。为在交互过程中保持模拟场景的对齐,我们开发了感知交互的场景同步方法,该方法可区分静态、握持和过渡模式。我们在四项真实机器人操作任务上评估了SpatialHarness,涵盖精确几何对齐、物体相对放置和关节物体交互。使用相同的冻结GPT-6 Astra策略,SpatialHarness大幅提升了任务成功率,其中插针任务从26.7%提升至66.7%,汉诺塔任务从0%提升至100%。这些结果表明,在测试时提升空间可观测性可释放强大多模态基础模型已具备的精细操作能力。项目网站:this https URL。
英文摘要
Frontier multimodal foundation models (e.g., GPT-6 Astra) have recently shown strong potential for direct robotic control, yet their performance on fine manipulation remains limited. We argue that an important source of failure is not necessarily insufficient policy capability, but insufficient spatial observability, where task-critical spatial relationships may be poorly revealed by the existing physical camera setup. We introduce SpatialHarness, a test-time embodied harness that provides test-time spatial scaffolding for fine robotic manipulation without policy fine-tuning or changes to the physical sensing setup. SpatialHarness maintains an online simulated scene synchronized with real-world execution, identifies task-critical spatial relationships, and renders complementary virtual views that expose them to a frozen multimodal policy. To keep the simulated scene aligned during interaction, we develop interaction-aware scene synchronization that distinguishes static, held, and transition modes. We evaluate SpatialHarness on four real-robot manipulation tasks spanning precise geometric alignment, object-relative placement, and articulated-object interaction. Using the same frozen GPT-6 Astra policy, SpatialHarness substantially improves task success, including from 26.7% to 66.7% on plug insertion and from 0% to 100% on Tower of Hanoi. These results indicate that improving spatial observability at test time can unlock fine-manipulation capabilities already present in strong multimodal foundation models. Project website: https://emilia113.github.io/SpatialHarness/.