从被动执行到主动探索:真实环境中的智能体具身操作
From Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments
- Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有具身操作框架被动执行、难以应对真实环境中语义线索与不可见目标的问题,提出由规划、感知、执行三模块协作的主动探索框架及感知-执行交错策略,在真实Find-and-Place任务中验证了有效性。
AI中文摘要:
近期智能体系统的进展显著增强了具身操作的长时程能力。然而,许多现有框架仍遵循被动执行范式,这限制了它们在涉及文本语义线索、干扰物以及初始不可见目标的真实场景中的适用性。为弥补这一差距,我们提出了一种基于智能体的主动探索框架,使机器人能够与环境动态交互,而非仅仅执行预定义指令。具体而言,我们的框架由三个协作模块组成:用于高层任务推理的规划模块、用于视觉场景理解的感知模块,以及用于低层操作的执行模块。这一设计使机器人能够主动获取任务相关信息,根据环境反馈调整自身行为,并在部分可观测条件下完成操作任务。此外,我们引入了一种细粒度的感知-执行交错策略,该策略将视觉反馈与技能执行紧密耦合,以提高探索的鲁棒性。我们在一个真实的Find-and-Place任务上评估了我们的方法,证明了其在目标物体必须在操作前被主动发现的挑战性环境中的有效性。
英文摘要:
Recent advances in agentic systems have substantially enhanced the long-horizon capability of embodied manipulation. However, many existing frameworks still follow a passive execution paradigm, which limits their applicability to real-world scenarios involving textual semantic cues, distractors, and initially invisible targets. To bridge this gap, we propose an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions. Specifically, our framework consists of three collaborative modules: a planning module for high-level task reasoning, a perception module for visual scene understanding, and an execution module for low-level manipulation. This design allows the robot to actively acquire task-relevant information, adapt its behavior based on environmental feedback, and complete manipulation tasks under partial observability. Furthermore, we introduce a fine-grained perception-execution interleaving strategy, which tightly couples visual feedback with skill execution to improve exploration robustness. We evaluate our method on a realistic Find-and-Place task, demonstrating its effectiveness in challenging environments where target objects must be actively discovered before manipulation.