CR-VLA-Force:学习控制感知的柔顺VLA模型用于鲁棒的接触丰富机器人操作
CR-VLA-Force: Learning Control-aware Compliance VLA Model for Robust Contact-rich Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
提出控制感知柔顺VLA框架,采用多模态专家混合与多阶段训练,结合自适应柔顺控制器,提升接触丰富操作中的力跟踪精度与成功率。
中文摘要 AI 辅助
将视觉运动策略或视觉-语言-动作(VLA)模型与力/力矩(F/T)感知相结合,已在机器人操作的模仿学习中展现出显著进展。然而,现有的力感知VLA模型在精确力跟踪和快速连续调整方面经常表现出有限的能力。这一缺陷源于动作块执行策略的局限性以及感知与实时控制之间的显著延迟。此类限制可能导致任务失败和安全风险,尤其是在执行动作块时施加过大的交互力而未能及时调整的情况下。为克服这一挑战,我们提出了控制感知柔顺VLA(CC-VLA)框架,用于反应式控制。CC-VLA模型采用多模态专家混合(MoE)架构来编码力信号序列和视觉-语言融合特征。此外,它利用多阶段训练策略,以确保在视觉-语义空间中的稳健感知以及稀疏采样条件下的有效力感知。另外,设计了一个VLA引导的自适应柔顺控制器,以在无接触运动中实现精确位置跟踪,并在接触丰富任务中实现最优力-位置跟踪。为促进高精度F/T数据采集,我们还实现了一种对抗性共享遥操作策略,用于接触丰富的演示,以增强系统安全性和交互性。大量真实世界实验表明,CC-VLA在具有挑战性的力感知任务中显著提高了成功率,并增强了力控制精度,同时在测试的部分分布外(partial-OOD)位姿偏移设置下提供了多层次的安全性和鲁棒性。
英文摘要
Integrating visuomotor policies or Vision-Language-Action (VLA) models with force/torque (F/T) perception has demonstrated significant progress in imitation learning for robotic manipulation. However, existing force-aware VLA models frequently exhibit limited capability in precise force tracking and rapid successive adjustments. This deficiency stems from the limitations of action-chunk execution strategies and the substantial latency between perception and real-time control. Such limitations can lead to task failures and safety risks, particularly when the execution of an action chunk exerts excessive interaction forces without timely adjustment. To overcome this challenge, we propose the Control-aware Compliance VLA (CC-VLA) framework for reactive control. The CC-VLA model employs a multimodal mixture-of-experts (MoE) to encode force signal sequences and vision-language fused feature. Furthermore, it utilizes a multi-stage training strategy to ensure robust perception within the visual-semantic space and effective force perception under sparse sampling conditions. Additionally, a VLA-guided adaptive compliance controller is designed to facilitate precise position tracking during contact-free motion and optimal force-position tracking for contact-rich tasks. To facilitate high-precision F/T data acquisition, we also implement an adversaria shared teleoperation strategy for contact-rich demonstrations that bolsters system safety and interactivity. Extensive real-world experiments demonstrate that CC-VLA significantly improves success rates in challenging force-perception tasks and enhances force-control precision, while providing multi-level safety and robustness under the tested partial-OOD pose-shift settings.
发表机构
- Huawei Technologies Co., Ltd.(华为技术有限公司)
- South China University of Technology(华南理工大学)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。