arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboHarness:用于长期规划的异构机器人策略的内存驱动编排

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning

Jinbang Huang, Zhiyuan Li, Yuanzhao Hu, Ran Qi, Yixin Xiao, Mark Coates, Tongtong Cao, Zhanguang Zhang, Yingxue Zhang

arXiv 2607.18060首次发表:更新:

发表机构

Huawei Noah’s Ark Lab; University of British Columbia; University of Toronto; McGill University; Labs(华为诺亚方舟实验室; 英属哥伦比亚大学; 多伦多大学; 麦吉尔大学; 2012实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长期机器人任务需多种能力、异构策略编排难的问题,提出RoboHarness框架,通过多模态执行内存等表征策略能力边界,经内存桥接稳定策略交接,实验验证其在长期规划和分布外鲁棒性上有显著提升。

AI 中文摘要

长期的机器人任务需要多种能力,单一策略无法可靠提供。异构策略具有互补优势,但编排它们需要处理不确定的能力边界和跨策略分布不匹配问题,现有基于同质、预定义技能且适用性固定的规划方法大多忽略了这些问题。我们提出了RoboHarness,一个统一框架,将独立开发的机器人控制系统封装为可重用的智能技能。它使用多模态执行内存和在线证据来表征策略能力边界以进行能力感知分解和路由。其内存桥接可稳定策略交接,通过检索与下一策略相关的执行轨迹等方式引导机器人。在多个基准测试、定制任务及真实机器人实验中验证了其有效性,在零样本长期规划和分布外鲁棒性方面有显著提升。

英文摘要

Long-horizon robotic tasks require a breadth of capabilities beyond what any single existing robot control policy can reliably provide. Combining heterogeneous policies with complementary strengths offers a promising solution, but introduces two key challenges: uncertain capability boundaries and distribution mismatches during policy handoffs. These challenges remain largely unaddressed by existing planning methods, which typically assume homogeneous, predefined skills with fixed applicability. We propose RoboHarness, a unified framework that encapsulates independently developed heterogeneous policies, including vision-language-action models (VLAs), world-action models (WAMs), reinforcement learning (RL) policies, and task and motion planners (TAMP), as reusable agentic skills. RoboHarness integrates understanding, memory, and evolution skills to reason about policy capabilities and support capability-aware task decomposition and policy routing. To mitigate distribution mismatches during policy handoffs, we introduce Memory Bridge, a plug-in policy-chaining mechanism that enables reliable transitions between heterogeneous policies without joint retraining. Extensive experiments across five public benchmarks, 500 customized tasks across 10 classes, and 135 real-robot trials demonstrate substantial gains in long-horizon and memory-dependent tasks, as well as robustness to out-of-distribution conditions.

CommentsBest Paper Award at ECCV 2026 Agent in the World Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑