发表机构
Harvard University; ETH Zürich(哈佛大学; 苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OCC4M提出以对象为中心的4D记忆,通过显式时空关系推理,在长时程操作中显著优于原始历史基线,实现高记忆与端到端成功率。
AI 中文摘要
长时程操作通常需要对当前视野中不存在的状态进行推理,例如消失物体的位置、时间身份或被打乱容器中的内容。我们提出OCC4M(“Occam”),一种以对象为中心的4D记忆,它在共享世界坐标系中维护持久轨迹,并显式表示时间、运动和包含关系。视觉语言模型(VLM)查询这种结构化记忆,以选择可操作的目标,用于无历史记录的低层执行。在七个模拟条件和350个回合中,OCC4M实现了96.6%的记忆成功率和88.9%的端到端成功率,而使用Gemini 3.7 Flash并具有完整观察历史和相同执行器的原始历史VLM基线FrameSamp分别仅为54.6%和57.7%。在受控的视点转移测试中,OCC4M在视点变化后保持100%的记忆成功率和98%的端到端成功率,而全历史FrameSamp则降至接近零的成功率。在20个固定相机Franka回合中,OCC4M达到85%的联合记忆准确率,而FrameSamp在从K=16到完整历史的上下文大小中最多仅为30%,并完成了45%的完整两阶段任务。这些结果支持显式的以对象为中心的记忆用于长时程操作中的持久时空推理。定性视频可在此https URL获取。
英文摘要
Long-horizon manipulation often requires reasoning about state absent from the current view, such as a vanished object's location, temporal identity, or the contents of a shuffled container. We present OCC4M ("Occam"), an object-centric 4D memory that maintains persistent tracks in a shared world frame and explicitly represents temporal, motion, and containment relations. A vision-language model (VLM) queries this structured memory to select actionable targets for history-free low-level execution. Across seven simulation conditions and 350 episodes, OCC4M achieves 96.6% memory success and 88.9% end-to-end success, versus 54.6% and 57.7% for FrameSamp, a raw-history VLM baseline using Gemini 3.7 Flash with the complete observation history and the same executor. In a controlled viewpoint-transfer test, OCC4M maintains 100% memory and 98% end-to-end success after a viewpoint change, while full-history FrameSamp falls to near-zero success. On 20 fixed-camera Franka episodes, OCC4M reaches 85% joint memory accuracy, versus at most 30% for FrameSamp across context sizes from $K=16$ to the complete history, and completes 45% of full two-stage tasks. These results support explicit object-centric memory for persistent spatiotemporal reasoning in long-horizon manipulation. Qualitative videos are available at https://occ4m-sup.github.io/occ4m-supplementary/.
Comments13 pages, 11 figures, 6 tables. Supplementary videos: https://occ4m-sup.github.io/occ4m-supplementary/