arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33197cs.ROcs.AI

TAO-DA:迈向自主操作——一种用于协调操作的双臂视觉-语言-动作模型

TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation

Yongsheng Zhao, Han Gao, Baoping Cheng, Jingyao Tang, Dian Zhou, Deng Liang, Ji Ge, Xuanzhang Wen, Lei Zhao, Ye Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有VLA模型在双臂操作中缺乏双臂状态解耦导致交叉干扰的问题,提出对称双臂专家架构与两阶段意图路由及任务进度预测模块,实现协调操作,并验证了技能泛化与迁移。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型为将高层语义信息 grounding 到低层机器人动作提供了统一框架,使得机器人能够在多种任务中进行可扩展的操作。然而,现有的 VLA 模型缺乏显式机制来解耦双臂的状态和意图,导致意外的双臂交叉干扰,从而降低任务执行成功率。为解决此问题,我们提出了一种对称的双臂专家(DAE)架构,该架构基于共享的视觉-语言模型(VLM)主干,并具有解耦的、特定于手臂的专家塔。专家选择通过两阶段双臂意图路由方案进行,其中专家在第一阶段由显式语言指令路由,或在第二阶段由隐式视觉语义路由。此外,我们引入了一个轻量级的任务进度预测模块,该模块利用预分块时间特征与本体感觉和视觉观察的语义表示之间的交叉注意力,来准确估计逐帧的任务完成进度。该模块促进任务进度同步,以支持协作多机器人任务的协调调度。实验结果表明,我们的模型在双臂意图路由和双臂交叉干扰解耦方面有效,并进一步提供了从单臂到双臂任务(以及反向)的技能泛化涌现的初步证据,以及跨臂运动域技能迁移。

英文摘要

Vision-Language-Action (VLA) models provide a unified framework for grounding high-level semantic information into low-level robot actions, enabling scalable robotic manipulation across diverse tasks. However, existing VLA models lack explicit mechanisms to disentangle the states and intents of the two arms, leading to unintended cross-arm interference that degrades task execution success. To address this issue, we propose a symmetric Dual-Arm Expert (DAE) architecture built upon a shared Vision-Language Model (VLM) backbone with decoupled, arm-specific expert towers. Expert selection is carried out through a two-stage dual-arm intent routing scheme, in which experts are routed either by explicit language instructions in the first stage or by implicit visual semantics in the second stage. Moreover, we introduce a lightweight task progress prediction module that leverages cross-attention between the pre-chunk temporal features and semantic representations of proprioceptive and visual observations to accurately estimate frame-wise task completion progress. This module facilitates task progress synchronization to support coordinated scheduling for collaborative multi-robot tasks. Experimental results demonstrate the effectiveness of our model in dual-arm intent routing and the disentanglement of cross-arm interference, and further provide preliminary evidence of emergent skill generalization from single- to dual-arm tasks (as well as the reverse), together with cross-arm motion-domain skill transfer.

↑