发表机构
Tongji University; AIRC, Midea Group(同济大学; 美的集团人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程机器人操作中多相同物体按序操作难的问题,提出MaskHarness-WAM框架,通过目标掩码连接高层规划与低层策略,并在子任务边界验证更新掩码,实验证明其显著优于有限时程策略。
AI 中文摘要
长时程机器人操作不仅需要稳定的局部视觉运动控制,还需要在整个执行过程中持续进行目标跟踪和可靠的任务进度评估。当多个物体外观完全相同且必须按指定顺序操作时,这一挑战尤为关键。在此类场景中,仅依赖有限时程的操作策略往往不足以确定应操作哪个实例以及任务何时应过渡到下一阶段。为解决这一问题,我们提出了MaskHarness-WAM,一种面向长时程操作的实例锚定驾驭框架。该系统通过目标掩码将高层任务规划与低层操作策略相连接,同时利用视觉反馈进行子任务调度和持续执行。由于每个子任务对应不同的目标实例,低层策略在每次子任务过渡时都需要在新场景下建立新的初始目标掩码。该驾驭框架在子任务边界持续重新观察环境,生成并验证目标掩码,从而更新提供给低层策略的实例级空间条件。此外,系统通过根据每个子任务的已验证完成状态切换目标实例来推进操作过程。在真实机器人平台上的实验表明,MaskHarness-WAM在顺序多物体操作上显著优于有限时程策略,展示了其在将局部操作技能扩展为可靠的长时程执行方面的有效性。
英文摘要
Long-horizon robot manipulation requires not only stable local visuomotor control, but also continuous target tracking and reliable task progress assessment throughout execution. This challenge becomes particularly critical when multiple objects share identical appearances and must be manipulated in a prescribed order. In such scenarios, relying solely on a limited-horizon manipulation policy is often insufficient to determine which instance should be operated on and when the task should transition to the next stage. To address this challenge, we propose MaskHarness-WAM, an instance-grounded harness for long-horizon manipulation. The proposed system connects high-level task planning with low-level manipulation policies through target masks, while leveraging visual feedback for subtask scheduling and continuous execution. Since each subtask corresponds to a different target instance, the low-level policy requires a newly established initial target mask under the updated scene at each subtask transition. The harness continuously re-observes the environment, generates, and verifies the target mask at subtask boundaries, thereby updating the instance-level spatial condition provided to the low-level policy. Furthermore, the system advances the manipulation process by switching target instances according to the verified completion status of each subtask. Experiments on a real robot platform demonstrate that MaskHarness-WAM substantially outperforms limited-horizon policies on sequential multi-object manipulation, showing its effectiveness in extending local manipulation skills to reliable long-horizon execution.