发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Opt2VLA提出力感知的VLA框架,通过显式力命令和轨迹优化监督,提升人形机器人接触丰富任务中的力调节准确性与稳定性。
AI 中文摘要
人形机器人被期望在日常环境中执行各种人类级别的任务,其中许多任务需要精确调节交互力。尽管最近的视觉-语言-动作(VLA)模型在语义规划和视觉运动控制方面显示出潜力,但现有的人形系统主要通过几何运动目标来表示动作,并依赖专注于运动跟踪的全身控制器,对交互力的显式推理或控制有限。这一局限性在接触丰富的任务中尤为突出,因为在这些任务中,几何上相似的运动可能根据任务上下文需要不同的力模式,并且视觉观察在接触后可能变得不可靠。在这项工作中,我们提出了Opt2VLA,一个力感知的VLA框架,在VLA到控制的接口处引入显式力命令,用于人形全身操作。一个单一的多任务VLA策略联合预测几何运动目标和连续的接触力参考,这些参考由基于任务特定强化学习(RL)的全身控制器跟踪。为了提供可扩展且物理上可靠的监督,我们通过带有显式力参考的全身轨迹优化(TO)生成动态可行且接触一致性的训练数据。我们在三个接触丰富的人形任务上评估了Opt2VLA,并表明显式力条件化能够比仅运动控制实现更准确和一致的力调节,而来自TO的物理可靠的扭矩监督进一步提高了力跟踪的准确性和稳定性。闭环评估进一步展示了在仿真和人形硬件上使用Opt2VLA进行语言条件化的力调制。
英文摘要
Humanoid robots are expected to perform diverse human-level tasks in daily environments, many of which require precise regulation of interaction forces. While recent vision-language-action (VLA) models have shown promise for semantic planning and visuomotor control, existing humanoid systems primarily represent actions through geometric motion goals and rely on whole-body controllers focused on motion tracking, with limited explicit reasoning or control of interaction forces. This limitation is particularly relevant in contact-rich tasks, where geometrically similar motions may require different force regimes depending on the task context and where visual observations may become unreliable after contact. In this work, we present Opt2VLA, a force-aware VLA framework that introduces explicit force commands at the VLA-to-control interface for humanoid whole-body manipulation. A single multi-task VLA policy jointly predicts both geometric motion goals and continuous contact-force references, which are tracked by task-specific reinforcement learning (RL)-based whole-body controllers. To provide scalable and physically grounded supervision, we generate dynamically feasible and contact-consistent training data via whole-body trajectory optimization (TO) with explicit force references. We evaluate Opt2VLA on three contact-rich humanoid tasks and show that explicit force conditioning enables more accurate and consistent force regulation than motion-only control, while physically grounded torque supervision from TO further improves force tracking accuracy and stability. Closed-loop evaluations further demonstrate language-conditioned force modulation with Opt2VLA in simulation and on humanoid hardware.