arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VT-Bridge:通过轻量级残差适配将预训练基础VLA桥接到VTLA

VT-Bridge: Bridging Pretrained Foundation VLAs to VTLAs via Lightweight Residual Adaptation

Yansong Wu, Tuo Yang, Rongping Zhao, Lingyun Chen, Xiao Chen, Junnan Li, Fan Wu, Alois Knoll

arXiv 2609.22606首次发表:更新:

发表机构

Technical University of Munich; Mohamed Bin Zayed University of Artificial Intelligence; Shanghai University(慕尼黑工业大学; 穆罕默德·本·扎耶德人工智能大学; 上海大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VT-Bridge提出轻量级残差适配策略,将预训练VLA模型桥接为VTLA模型,仅需少量演示和98万参数适配器,将任务完成率从11.7%提升至62.9%。

AI 中文摘要

视觉-触觉-语言-动作(VTLA)模型在接触丰富的操作任务中已展现出优于视觉-语言-动作(VLA)模型的明显优势。然而,开发VTLA模型受到大量视觉-触觉数据和计算资源需求的严重制约。为解决这一瓶颈,我们提出了VT-Bridge,一种轻量级残差适配策略,将预训练的基础VLA桥接到VTLA。VT-Bridge并非从头训练VTLA模型或修改预训练VLA的原始架构,而是在不同VLA骨干网络上采用相同的轻量级残差适配器架构,并使用骨干网络特定的权重以机器人执行频率来细化动作。这一设计显著降低了数据和训练门槛。具体而言,每个任务仅需最多50个视觉-触觉演示即可微调VLA骨干网络,并训练一个拥有98万参数的残差适配器。在三个具有代表性的VLA骨干网络(π₀、π₀.₅和SmolVLA)上,跨四个接触丰富的操作任务进行的实验进一步证明了其在不同VLA架构上的一致有效性。平均而言,VT-Bridge将任务完成率从仅使用任务级VLA微调时的11.7%提升至62.9%。综上所述,这些发现证明了VT-Bridge在接触丰富操作中的广泛适用性、有效性和易用性。项目页面可在此https URL获取。

英文摘要

Vision-Tactile-Language-Action (VTLA) models have demonstrated clear advantages over Vision-Language-Action (VLA) models in contact-rich manipulation. However, developing VTLA models is severely constrained by the massive amounts of vision-tactile data and computational resources required. To address this bottleneck, we propose VT-Bridge, a lightweight residual adaptation strategy that bridges pretrained foundation VLAs to VTLAs. Rather than training a VTLA model from scratch or modifying the original architecture of a pretrained VLA, VT-Bridge employs an identical lightweight residual-adapter architecture across VLA backbones and uses backbone-specific weights to refine actions at the robot execution frequency. This design substantially lowers the data and training barriers. Specifically, it requires up to 50 vision-tactile demonstrations per task to fine-tune a VLA backbone and train a 0.98M-parameter residual adapter. Experiments with three representative VLA backbones ($π_0$, $π_{0.5}$, and SmolVLA) across four contact-rich manipulation tasks further demonstrate its consistent effectiveness across VLA architectures. On average, VT-Bridge raises the task completion rate from 11.7% with task-level VLA fine-tuning alone to 62.9%. Together, these findings demonstrate the broad applicability, effectiveness, and accessibility of VT-Bridge for contact-rich manipulation. The project page is available at https://hoxnocha.github.io/vt-bridge-web/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑