arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LIFD:用于机器人操作中三维感知场景记忆的锚定扩散

LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation

Wenbo Li, Yiteng Chen, Wenhao Li, Qingyao Wu

arXiv 2609.19796首次发表:更新:

发表机构

South China University of Technology(华南理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LIFD通过锚定扩散与多视图一致性学习三维感知场景记忆,从单视图和循环记忆完成场景表示,提升机器人操作成功率,在LIBERO和MetaWorld上分别达91.6%和79.8%。

AI 中文摘要

在部分可观测条件下的机器人操作需要超出当前视野的空间信息。几何感知的RGB特征描述了可见结构,但随着机器人或场景的移动,先前观察到的区域可能消失。因此,维持有用的场景表示需要在推断缺失内容的同时保留观察历史,且不丢失与可见证据的联系。我们引入了LIFD(观察、想象、聚焦与执行),一个用于持久、三维感知场景记忆的框架。LIFD从多视图一致性中学习场景令牌表示,并从单一RGB视图和循环记忆中完成该表示。一个修正流模型生成令牌,而锚定引导的交叉注意力将完成过程条件化于当前几何特征。紧凑的槽位特征将该表示连接到操作策略。表示学习期间使用多视图和几何监督;部署时仅需一个RGB相机、本体感觉和任务指令。LIFD(分阶段)在LIBERO上达到91.6%的平均成功率,在MetaWorld上达到79.8%,相比联合训练在LIBERO上的平均成功率提高了3.1个百分点。在四个UR5e任务族(每个族十个演示)上,它实现了56.0%的平均成功率,而OpenVLA-7B为40.5%。

英文摘要

During manipulation, robot and scene motion can move previously observed regions outside the camera's field of view. Geometry-aware RGB features encode visible structure, while control under partial observability requires scene memory that integrates observation history and grounds inferred content in current evidence. We introduce \lifd{} (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory. LIFD learns scene tokens through multi-view agreement, then completes them from a single RGB view and recurrent memory using rectified flow. Anchor-Guided Cross-Attention anchors generation to current geometry-aware features, and compact slot features condition a visuomotor policy. Multi-view and geometric supervision are used during representation learning; deployment requires one RGB camera, proprioception, and a task instruction. LIFD (Staged) reaches 91.6\% average success on LIBERO and 79.8\% on MetaWorld, improving LIBERO average success by 11.1 percentage points over Joint training. After policy-head adaptation with ten demonstrations per family, LIFD achieves 56.0\% mean success across four UR5e task families, compared with 40.5\% for OpenVLA-7B.

Comments8 pages, 4 figures. Submitted to ICRA 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑