ACE:通过零样本工作流推理实现具身操纵的智能体控制
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning
浏览论文内容
中文总结 AI 辅助
研究开放式桌面操作,提出ACE零样本工作流推理框架,结合智能体工作流推理与视觉基础接口、可重复使用的拾取和放置原语等技能,经掩码介导视觉动作接口连接语义推理与物理控制,能在线适应多种情况,在复杂任务中表现出色。
中文摘要 AI 辅助
开放式桌面操作要求智能体不仅理解自然语言,还能适应动态环境和执行失败。我们提出了ACE(具身操纵的智能体控制),这是一个用于从自然语言进行桌面拾取和放置的零样本工作流推理框架。ACE将智能体工作流推理与两种面向机器人的可执行技能相结合,通过掩码介导的视觉动作接口连接语义推理和物理控制,在多时间尺度内存支持的闭环中运行,能在线适应多种情况,在复杂任务中表现出色。
英文摘要
General-purpose manipulation requires both semantic reasoning over task constraints and reliable execution of contact-rich actions. We present ACE, an agentic manipulation harness that composes a high-level language agent with a reusable mask-conditioned visuomotor policy. Given an open-ended instruction, the agent solves semantic constraints, binds objects to destination roles, and decomposes the task into executable transfers represented by tracked pick-and-place masks. Execution feedback supports outcome assessment, re-grounding, and retry, while persistent object and task context preserves earlier associations when manipulation changes visible cues. We evaluate ACE on two physical multi-step tabletop tasks, Semantic Formula Assembly and Constraint Retrieval. The visuomotor policy is trained only on generic pick-and-place demonstrations and reused without complete demonstrations of either evaluation task, enabling task-level zero-shot composition. Across 20 randomized trials per task, ACE achieves 70% and 80% success, respectively, compared with 55% and 70% without persistent context. These results suggest that an agentic harness can extend a primitive-trained manipulation policy to semantically distinct tasks through explicit object-destination interfaces and closed-loop execution feedback.
发表机构
- Department of Computer Science, Tsinghua University(清华大学计算机科学系)
- National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
- Wuxi Dexteroushands Robotic Technology Co.(无锡灵犀机器人技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。