arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29166cs.ROcs.AI

HarnessPAI:物理AI的演进式驾驭框架

HarnessPAI: An Evolving Harness for Physical AI

Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan, Yang Li, Qing Li, Shangding Gu, Huichi Zhou, Shuqing Shi, Fei Ni, Shuo Lu, Weicheng Meng, Kan… 展开作者

Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan, Yang Li, Qing Li, Shangding Gu, Huichi Zhou, Shuqing Shi, Fei Ni, Shuo Lu, Weicheng Meng, Kang Li, Jin Wu, Kang Zhao, Shangmin Guo, Gen Li, Yongqiang Tang, Zhizhong Zhang, Yuan Xie, Heng Qu

首次发表
浏览论文内容

中文总结 AI 辅助

HarnessPAI提出一种与模型和具身方式无关的演进式驾驭框架,以代码为接口组织行为原语,通过开环执行与闭环演进分离时间尺度,在不重训模型的情况下显著提升物理AI任务表现,并可作为专家数据收集器。

中文摘要 AI 辅助

物理AI旨在构建具身智能体,使其能够感知世界、理解并推理世界,并决定如何行动。然而,该领域主要聚焦于最后一个组成部分:将观测映射到低级控制的行为模型。当前主流的训练方法可能会削弱实现稳健行为所需的感知和推理能力,即使强大的行为模型也容易受到场景扰动和长时程任务的影响。我们提出了HarnessPAI,一个与模型和具身方式无关的物理AI驾驭框架,将代码视为可执行且可演进的接口,用于组织底层行为原语。该框架区分了两个时间尺度:在一次滚动执行中,它在程序层面进行开环执行,由固定程序指导和检查执行过程;在多次滚动之间,它进行闭环演进,利用执行反馈来修订程序,并将失败经验提炼为可复用的技能。在桌面机械臂、家用机器人、扫地机器人和四足行走智能体上,HarnessPAI在无需重新训练底层模型的情况下,相较于纯行为模型和代码即策略基线均有提升:在LIBERO-PRO上比π_{0.5}高出61.6个百分点,在RoboCasa原子任务上比WorldDreamer高出27.2个百分点。一旦程序被选定,滚动执行无需在线高层LLM推理。除执行外,收敛后的程序还是一个廉价且可靠的专家数据收集器,在收集的专家数据上微调π_{0.5}可将LIBERO-PRO上的成功率提升38.8个百分点。我们的结果表明,物理AI的前沿不仅依赖于更强的行为模型,还依赖于将感知、任务理解与推理以及行为执行整合到一个统一、可验证且反馈驱动的系统中的可执行驾驭框架。网站:此https URL

英文摘要

Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. Yet the field has focused primarily on the last component: the action model that maps observations to low-level controls. The prevailing training recipe can erode the perceptual and reasoning capabilities needed for robust behavior, leaving even strong action models vulnerable to scene perturbations and long-horizon tasks. We introduce HarnessPAI, a model- and embodiment-agnostic Harness framework for Physical AI that treats code as the executable and evolvable interface that organizes the underlying action primitive. The framework separates two timescales: within a rollout, it executes open-loop at the program level, with a fixed program guiding and checking execution; across rollouts, it evolves closed-loop, using execution feedback to revise the program and distill failures into reusable skills. Across desktop robot arms, household robots, a robot vacuum, and a legged walking agent, HarnessPAI improves on both pure action models and code-as-policy baselines without retraining the underlying model: a 61.6-point gain over $π_{0.5}$ on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks. Once a program is selected, rollout execution requires no online high-level LLM deliberation. Beyond execution, the converged program is also a cheap and reliable expert-data collector, and fine-tuning $π_{0.5}$ on collected expert data lifts success rate on LIBERO-PRO by 38.8 points. Our results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system. Website: https://darwin-agent.github.io/HarnessPAI

补充信息

↑