arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03889cs.ROcs.AI

FWBC-VLA:面向接触丰富型移动操纵的力感知全身补偿

FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation

Yutian Zhang, Siyuan Ma, Liwen Yang, Yang Li, Ce Hao, Haozhen Chi, Dong Wei, Qiaojun Yu, Dibo Hou

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对现有VLA模型无法解释物理交互、WBC策略无法区分任务相关力与扰动的问题,提出FWBC-VLA力感知框架,结合无传感器接触估计与补偿控制,在轮腿机器人的接触丰富型移动操纵任务中验证了有效性。

中文摘要 AI 辅助

接触丰富型移动操纵需要在语义动作生成与物理交互控制之间建立桥梁。现有视觉-语言-动作(Vision-language-action, VLA)模型可从视觉和语言观测中生成任务级动作,但无法解释这些动作引发的物理交互。虽然全身控制(Whole-Body Control, WBC)策略可稳定机器人,但无法区分操纵过程中与任务相关的交互力和外部扰动引发的力。力/力矩传感器虽能提供物理交互的直接测量值,但加装这些传感器会产生额外硬件成本和大量集成工作,尤其对于未设计传感器集成的平台。为解决该问题,我们提出FWBC-VLA,一种面向轮腿式机器人的力感知框架,用于连接任务级VLA动作生成与低级全身补偿控制。首先,我们引入HSR-Force,一种无传感器残差力矩估计器,用于推断接触强度及其时间变化。这些接触估计值随后被编码为 token 并注入VLA动作专家的动作解码过程中,使策略能感知接触的开始、持续加载及释放。对于移动操纵任务,我们在包含超5000个回合的WL&Arm数据集上对预训练VLA主干的所有参数进行微调。此外,机器人的本体感受状态、雅可比矩阵推导的体坐标系力估计值及估计的接触状态会共同输入补偿生成器以生成校正动作。以操纵为中心的动作随后与校正动作结合并传递给WBC策略执行。在白板擦拭和带闭门器的开门任务上进行的现实实验证明了FWBC-VLA在接触丰富型移动操纵中的有效性。

英文摘要

Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction control. Existing Vision-language-action (VLA) models generate task-level actions from visual and linguistic observations, but cannot interpret the physical interactions induced by those actions. While the whole-body control (WBC) policy can stabilize the robot, it cannot distinguish task-relevant interaction forces from forces induced by external disturbances during manipulation. Although force/torque sensors provide direct measurements of physical interactions, retrofitting them entails additional hardware costs and substantial integration effort, particularly for platforms not designed with sensor integration in mind. To address this problem, we propose FWBC-VLA, a force-aware framework that bridges task-level VLA action generation and low-level whole-body compensation control for wheeled-legged robots. First, we introduce HSR-Force, a sensorless residual-torque estimator for inferring contact strength and its temporal variation. These contact estimates are then encoded as tokens and injected into the VLA action expert during action decoding, enabling the policy to perceive contact onset, sustained loading, and release. For loco-manipulation tasks, all parameters of the pretrained VLA backbone are fine-tuned on our WL\&Arm Dataset, which comprises more than 5,000 episodes. Moreover, the robot's proprioceptive state, the Jacobian-derived body-frame force estimate, and the estimated contact state are jointly fed into a compensation generator to produce corrective actions. The manipulation-centric actions are subsequently combined with the corrective actions and passed to the WBC policy for execution. Real-world experiments on whiteboard wiping and door opening with a door closer demonstrate the effectiveness of our FWBC-VLA in contact-rich loco-manipulation.

发表机构

  • Zhejiang University(浙江大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Tsinghua University(清华大学)
  • Zhongguancun Academy(中关村学院)
  • Deep Robotics(深度机器人公司)
  • Zhejiang University of Science and Technology(浙江科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑