具身场景重排规划
Embodied Scene Rearrangement Planning
- School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出具身场景重排规划(ESRP)新任务,构建含5400+场景对的ESRP-Bench基准,评估四类方法发现现有方法难高效完成,为智能体现实部署奠定基础。
AI中文摘要:
本文介绍了具身场景重排规划(Embodied Scene Rearrangement Planning,ESRP),这是一项新任务,要求具身智能体仅使用自我中心观测和自上而下的目标布局,在3D场景中重新排列家具以匹配目标配置。与先前的重排任务不同,ESRP不允许访问全局状态,还引入了物体间的相互遮挡,反映了现实世界机器人部署的实际约束。这些因素使得将部分自我中心观测与全局目标布局对齐,对于长程规划而言极具挑战性。为推动相关研究,我们推出了基于OmniGibson构建的综合基准ESRP-Bench,包含超过5400个场景对和8200个物体。我们定义了三个多级指标来评估重排质量,并提供了四个基线方法:一种分层任务与运动规划方法、一种基于视觉语言模型的方法,以及两种基于学习的方法(模仿学习IL和强化学习RL)。实验结果表明,当前方法难以高效完成该任务,凸显ESRP是具身智能体在场景理解和长程任务规划领域的一个极具挑战性的前沿方向。本研究为将智能体部署到现实场景中奠定了基础。项目页面:this https URL。
英文摘要:
This paper introduces Embodied Scene Rearrangement Planning (ESRP), a novel task requiring embodied agents to rearrange furniture in 3D scenes to match a target configuration using only egocentric observations and a top-down target layout. Unlike prior rearrangement tasks, ESRP precludes global state access and introduces mutual object occlusions, reflecting the practical constraints of real-world robotic deployment. These factors make aligning partial egocentric observations with the global target layout particularly challenging for long-horizon planning. To facilitate research, we present ESRP-Bench, a comprehensive benchmark built on OmniGibson featuring over 5,400 scene pairs and 8,200 objects. We define three multi-level metrics to evaluate rearrangement quality and provide four baselines: a hierarchical task-and-motion planning method, a vision-language-model-based method, and two learning-based approaches (IL and RL). Experimental results demonstrate that current methods struggle to complete the task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning. This work serves as a stepping stone toward deploying intelligent agents in real-world scenarios. Project page: https://pie-lab.cn/ESRP/.