arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40306cs.ROcs.AIcs.LG

DynaHarness:用于自进化机器人代理的动态物理约束装置

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

  • Nanyang Technological University(南洋理工大学)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang, Langning Yan, Hongye Cao, Ziwei Wang

AI总结:

针对长时程操作中语义与物理执行脱节的问题,提出动态物理约束装置DynaHarness,通过共享执行契约和失败归因实现能力自进化,在LIBERO-Pro上显著提升成功率。

AI中文摘要:

预训练的机器人策略提供了有用的动作先验,但长时程操作仍然需要语义推理与物理执行之间的协调。语义推理的时间尺度比物理交互更粗,而回合级失败只能提供有限的指导,说明应修改哪个系统组件。我们提出DynaHarness,一种动态物理约束装置,通过共享执行契约将语义推理与物理治理耦合,并将失败证据转化为经过验证的能力修订。具体而言,慢速大脑提出能力和符号参数,而快速大脑则对命令进行落地和监控,拒绝未解决的行动,替代能力,并在需要时请求重新规划。物理执行契约约束每个接受的命令,并记录跨分析技能、恢复技能和冻结VLA的执行证据。失败归因在这些记录中定位故障,并指导对可重用能力或执行机制进行有针对性的修订。成对的回归检查控制接纳或拒绝,从而闭环自进化循环。在LIBERO-Pro上,DynaHarness在800个新采样的初始状态下达到75.2%的成功率,而冻结策略仅为17.5%。使用相同的能力库,完整动态执行达到74.0%,而名义上的一步重规划为63.9%。这证明了DynaHarness作为一种动态物理约束装置的价值,它管理现有能力在执行过程中的落地、监控和协调。我们的项目页面位于此https URL。

英文摘要:

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised. We propose DynaHarness, a dynamic physical harness that couples semantic reasoning with physical governance through a shared execution contract and turns failure evidence into validated capability revisions. To be more specific, the slow brain proposes capabilities and symbolic arguments, while the fast brain grounds and monitors commands, refuses unresolved actions, substitutes capabilities, and requests replans when needed. The physical execution contract bounds each accepted command and records execution evidence across analytic skills, recovery skills, and the frozen VLA. Failure attribution localizes faults in these records and directs targeted revisions of reusable capabilities or execution mechanisms. Paired regression checks govern admission or rejection, closing the self-evolution loop. On LIBERO-Pro, DynaHarness achieves 75.2% on 800 newly sampled initial states, compared with 17.5% for the frozen policy. With the same capability library, full dynamic execution reaches 74.0% versus 63.9% under nominal one-step replanning. This demonstrates the value of DynaHarness as a dynamic physical harness that governs how existing capabilities are grounded, monitored, and coordinated during execution. Our project page is at https://denghaoyuan123.github.io/Dynaharness_page/.

补充信息

↑