$N_0$-VTLA:使用潜在触觉令牌扩展视觉-触觉-语言-动作模型
$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- NeoteAI Team(NeoteAI团队)
- Fudan TEAI Team(复旦大学TEAI团队)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究提出$N_0$-VTLA模型,基于视觉主干,通过视觉-触觉预训练、分阶段触觉路径集成和优势条件离线策略改进进行触觉集成训练。该模型在丰富接触基准测试中表现出色,为通用触觉驱动操作策略奠定基础。
AI中文摘要:
我们提出了$N_0$-VTLA,一个视觉-触觉-语言-动作(VTLA)基础模型,它能够通过触觉感知和触觉反馈控制进行细粒度的丰富接触操作,并能从存储的部署数据中进行离线策略改进。基于当前基于视觉的主干,我们提出了一种触觉集成训练方法,包括视觉-触觉预训练、分阶段触觉路径集成和优势条件离线策略改进。预训练期间,策略从我们的大规模视觉-触觉机器人数据集NeoData中学习广泛的接触先验;$N_0$-VTLA是首个大规模在触觉数据上预训练的VTLA模型。训练后,我们用预测触觉路径增强策略,将大规模学习的接触模式提炼为下游以触觉为中心的操作所需的精细运动调整。对于离线策略改进,我们引入ALTER,一种优势条件离线强化学习方法,将相对进展和轨迹事件比较转换为二进制优势标签,用于在固定部署语料库上进行策略训练,进一步改进了对丰富接触技能(如可变形物体操作)的特定任务学习。在丰富接触基准测试中,$N_0$-VTLA大幅优于强大的基线:它赢得了所有九个真实机器人NeoReal任务,并在二十任务模拟套件上达到63.8%的平均成功率,而最强基线为44.0%。用ALTER训练的$N_0$-VTLA策略在三个长时真实机器人任务上成功率达到75-95%。这些结果为通用的触觉驱动操作策略奠定了基础。
英文摘要:
We present $N_0$-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, $N_0$-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, $N_0$-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. $N_0$-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.