arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14310cs.RO

VLBiMan++:扩展视觉-语言锚定的一次性双臂操作泛化边界

VLBiMan++: Expanding the Generalization Boundary of Vision-Language Anchored One-Shot Bimanual Manipulation

  • Shenzhen University(深圳大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • DexForce
  • Imperial College London(伦敦帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Huayi Zhou, Wei Gao, Yiyang Han, Kui Jia, Hui Huang

AI总结:

VLBiMan++基于单次人类演示,通过任务分解和视觉-语言几何适应,沿任务、物体、场景、本体和部署五维度扩展双臂操作泛化,实验验证其强适应性和成功率。

AI中文摘要:

可泛化的双臂机器人操作需要一个可复用的任务先验,该先验能够持续适用于日益多样化的任务、物体、场景、本体和执行条件,从而避免大规模遥操作演示和策略重新训练的高昂成本。在本工作中,我们提出了VLBiMan++,一个扩展框架,用于扩展视觉-语言锚定的一次性双臂操作的泛化边界。从单个人类演示出发,VLBiMan++执行任务感知分解以识别可复用且可适应的技能组件,并采用视觉-语言引导的几何适应将这些技能迁移到新配置中而无需重新训练。在此基础之上,我们系统地沿五个维度扩展泛化:通过多样化和长时程技能组合实现任务泛化;跨未见类别、不同几何形状以及更复杂的铰接或可变形物体的物体泛化;在杂乱、遮挡和动态干扰下的场景泛化;跨异构双臂机器人平台的本体泛化;以及在重复外部扰动下通过长时间闭环执行实现的部署泛化。为支持这一更广泛的范围,我们进一步引入了物体状态感知适应和轻量级轨迹优化机制,以适应超出简单刚性6自由度位姿变化的情况,同时保持可靠的双臂协调。大量真实世界实验表明,VLBiMan++在这些日益具有挑战性的设置中保持了强大的任务成功率和适应能力。总体而言,VLBiMan++将一次性双臂操作从展示孤立的可迁移性推进到更系统化和可扩展的框架,用于跨任务、物体、场景、本体和长期部署条件的泛化。

英文摘要:

Generalizable bimanual robotic manipulation requires a reusable task prior that can persist across increasingly diverse tasks, objects, scenes, embodiments, and execution conditions, thus avoiding the prohibitive cost of large-scale teleoperated demonstrations and policy retraining. In this work, we present VLBiMan++, an extended framework that expands the generalization boundary of vision-language anchored one-shot bimanual manipulation. Starting from a single human demonstration, VLBiMan++ performs task-aware decomposition to identify reusable and adaptable skill components, and employs vision-language grounded geometric adaptation to transfer these skills to novel configurations without retraining. Building on this foundation, we systematically extend generalization along five dimensions: task generalization through diverse and long-horizon skill compositions; object generalization across unseen categories, varying geometries, and more complex articulated or deformable objects; scene generalization under clutter, occlusion, and dynamic interference; embodiment generalization across heterogeneous dual-arm robotic platforms; and deployment generalization through prolonged closed-loop execution under repeated external perturbations. To support this broader scope, we further introduce object-state-aware adaptation and lightweight trajectory optimization mechanisms that accommodate changes beyond simple rigid 6-DoF pose variations while preserving reliable bimanual coordination. Extensive real-world experiments demonstrate that VLBiMan++ maintains strong task success and adaptation capability across these increasingly challenging settings. Overall, VLBiMan++ advances one-shot bimanual manipulation from demonstrating isolated transferability toward a more systematic and scalable framework for generalization across tasks, objects, scenes, embodiments, and long-term deployment conditions.

补充信息

↑