CompVLA:用于接触丰富操作的可变柔顺视觉-语言-动作模型
CompVLA: A Variable Compliance Vision-Language-Action Model for Contact-rich Manipulation
浏览论文内容
中文总结 AI 辅助
针对现有VLA模型缺乏柔顺控制导致接触丰富任务性能差的问题,提出联合预测运动与刚度矩阵的CompVLA框架,通过柔顺专家模块和阻抗控制提升成功率,实验验证其有效性。
中文摘要 AI 辅助
接触丰富操作要求机器人不仅调节运动,还要调节其对外力的屈服方式,已成为视觉-语言-动作(VLA)模型的下一个前沿领域。然而,现有VLA仅输出纯运动学命令,在真实世界的接触丰富任务中性能下降。本文提出CompVLA,一个统一的VLA框架,从RGB图像和语言输入中联合预测运动与刚度矩阵。我们的方法通过专用柔顺专家模块增强传统架构,该模块输出时变刚度和虚拟位移轮廓,并通过几何阻抗控制执行。我们证明CompVLA在多种接触丰富任务中取得最高平均成功率,优于普通和柔顺感知VLA基线,消融实验确认每个组件均不可或缺。
英文摘要
Contact-rich manipulation, requiring robots to regulate not only motion but also how they yield to external forces, has emerged as the next frontier for Vision-Language-Action (VLA) models. However, existing VLAs output purely kinematic commands, degrading performance on real-world contact-rich tasks. In this paper, we introduce CompVLA, a unified VLA framework that jointly predicts motion and stiffness matrix from RGB and language inputs. Our approach augments the conventional architecture with a dedicated Compliance Expert, which outputs time-varying stiffness and virtual displacement profiles executed via geometric impedance control. We demonstrate that CompVLA achieves the highest average success rate across diverse contact-rich tasks, outperforming both vanilla and compliance-aware VLA baselines, with ablations confirming each component is essential.
发表机构
- Seoul National University(首尔大学)
- Microsoft Research Asia(微软亚洲研究院)
机构由 AI 辅助整理,请以论文原文为准。