TAO-Force:面向接触丰富操作任务的力感知与快慢控制统一框架
TAO-Force: Unifying Force-Aware Perception and Fast-Slow Control for Contact-Rich Manipulation
- China Mobile (Hangzhou) Information Technology Co., Ltd.(中国移动(杭州)信息技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TAO-Force提出力条件化VLA框架,通过F-FiLM融合力感知与接触门控快慢控制,解决接触丰富操作中视觉感知不足和位置控制不柔顺的问题,实验验证其有效性与鲁棒性。
AI中文摘要:
视觉-语言-动作(VLA)模型已在多种机器人操作任务中展现出强大的性能,但其以视觉为主的感知方式和位置控制的执行模式在接触丰富的操作任务中仍显不足。仅凭视觉观察往往难以提供接触起始和交互强度的充分证据,而位置控制策略无法对快速变化的接触动力学做出柔顺响应。为同时弥合感知与控制两方面的不足,我们提出了TAO-Force,一种力条件化的VLA框架,将力感知策略学习与接触调节执行相结合。在力感知方面,TAO-Force引入了力条件化特征级线性调制(F-FiLM),将编码后的力反馈注入冻结的预训练视觉-语言骨干网络的表示中,同时保留其语义先验。在响应控制方面,它采用接触门控的快慢架构,其中慢速位置控制分支在非接触阶段跟踪标称轨迹,快速导纳控制分支在接触阶段调节物理交互。在力感知任务上的详细分析以及在四个接触丰富操作任务上的真实世界评估验证了TAO-Force的有效性和鲁棒性。
英文摘要:
Vision-Language-Action (VLA) models have demonstrated strong performance across diverse robotic manipulation tasks, yet their predominantly vision-centric perception and position-controlled execution remain insufficient for contact-rich manipulation. Visual observations alone often provide limited evidence of contact onset and interaction magnitude, while position-control policies cannot respond compliantly to rapidly changing contact dynamics. To bridge both the perception and control gaps, we propose TAO-Force, a force-conditioned VLA framework that combines force-aware policy learning with contact-regulated execution. For force-aware perception, TAO-Force introduces Force-conditioned Feature-wise Linear Modulation (F-FiLM) to inject encoded force feedback into the representations of a frozen pretrained visual-language backbone while preserving its semantic priors. For responsive control, it employs a contact-gated fast-slow architecture, with a slow position-control branch tracking nominal trajectories during non-contact phases and a fast admittance-control branch regulating physical interaction during contact phases. Detailed analyses on a force-perception task and real-world evaluations across four contact-rich manipulation tasks validate the effectiveness and robustness of TAO-Force.