发表机构
The University of Hong Kong; Shanghai Jiao Tong University; Xense Robotics; University of Pennsylvania(香港大学; 上海交通大学; Xense Robotics; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PEARS提出物理先验引导的混合强化学习框架,通过失败推理与扩散引导,在触觉操作中实现样本高效在线适配,显著提升成功率并减少交互次数。
AI 中文摘要
预训练的机器人策略在部署过程中遇到分布外(OOD)条件时,可能会遭受显著的性能下降,这促使通过真实世界交互进行后训练。然而,基于强化学习(RL)的后训练通常需要大量的环境交互,这一负担在操作任务中尤为突出,因为每次试验可能缓慢、昂贵或具有破坏性。因此,我们提出了PEARS,一种物理先验引导的混合强化学习框架,用于利用触觉反馈对预训练策略进行样本高效的在线适配。在每个回合之后,其物理引导的力推理(PFR)模块利用编码在视觉语言模型(VLM)中的物理先验,从视觉结果和触觉交互历史中诊断失败,并更新任务适当的接触力界限。一个高频混合力-位置控制器随后在接触期间强制执行这些界限。补充地,触觉条件扩散引导强化学习调整冻结的流匹配策略的潜在噪声,以纠正自由空间运动和接触时间中的错误,而无需更新基础模型。在仿真中,PEARS相比最强的每任务基线,成功率提高了12.4至37.4个百分点。PEARS还将达到特定成功阈值所需的交互回合数相对于最快基线减少了最多53.2%。在真实世界实验中,PEARS在擦白板任务中达到了95%的成功率,在移液管液体吸取任务中达到了90%的成功率。这些结果表明,将PFR模块与策略引导相结合可以加速适配,同时减少昂贵的交互。项目网站可在以下URL访问:this https URL。
英文摘要
Pretrained robotic policies can suffer substantial performance degradation under out-of-distribution (OOD) conditions encountered during deployment, motivating post-training through real-world interaction. However, reinforcement-learning (RL)-based post-training typically requires substantial environment interactions, a burden that is especially significant in manipulation, where each trial can be slow, costly, or destructive. Therefore, we present PEARS, a physics-prior-guided hybrid RL framework for sample-efficient online adaptation of pretrained policies with tactile feedback. After each episode, its physics-guided force reasoning (PFR) module uses physical priors encoded in a vision-language model (VLM) to diagnose failures from the visual outcome and tactile interaction history and update task-appropriate contact-force bounds. A high-frequency hybrid force-position controller then enforces these bounds during contact. Complementarily, tactile-conditioned diffusion steering reinforcement learning adjusts the latent noise of the frozen flow-matching policy to correct errors in free-space motion and contact timing without updating the base model. In simulation, PEARS improves success rates by 12.4-37.4 percentage points over the strongest per-task baselines. PEARS also reduces the number of interaction episodes required for a certain success threshold by up to 53.2% relative to the fastest baseline. In real-world experiments, PEARS achieves success rates of 95% on Whiteboard Erasing and 90% on Pipette Liquid Aspiration. These results show that combining the PFR module with policy steering can accelerate adaptation while reducing costly interactions. The project website is available at https://song-kun.github.io/pears.
Comments9 pages, 4 figures. Project website: https://song-kun.github.io/pears