AI 中文总结
研究旨在解决学习触觉表示并应用于VLA模型的挑战,提出τ框架从未来视觉监督学习动作条件时空触觉表示,结合视觉语言特征用于动作生成,引入TacAura数据集,实验证明其性能优于现有模型且能泛化。
AI 中文摘要
在数据和建模层面,学习信息丰富的触觉表示并有效将其应用于预训练的视觉-语言-动作(VLA)模型仍具有挑战性。数据层面,特定任务演示有限限制表示质量,大规模预训练成本高;建模层面,现有方法或关注瞬时接触状态,或用6D扳手序列建模时间交互动态,高维触觉信号未充分探索。为应对这些挑战,我们提出τ,一个触觉增强的VLA框架,从未来视觉监督中学习动作条件时空触觉表示,并在有限数据下与视觉语言特征融合用于动作生成。我们还引入TacAura数据集。实验表明τ优于现有模型并能推广到未见物体和场景,提升了操作性能和鲁棒性。
英文摘要
Incorporating tactile sensing into Vision-Language-Action (VLA) models holds promise for contact-rich manipulation, where visual observations alone often fail to capture critical cues about physical interactions. However, learning informative tactile representation while effectively adapting it to pretrained VLA models remains challenging under limited task-specific data. Existing methods either focus on instantaneous contact states or model temporal interaction dynamics using 6D wrench sequences, leaving high-dimensional tactile signals underexplored. To address these challenges, we present τ, a touch-augmented VLA framework that learns an action-conditioned spatiotemporal tactile representation from future visual supervision inspired by the Joint-Embedding Predictive Architecture (JEPA), and fuses it with vision-language features for action generation. This supervision operates in latent space and is used only during training, adding no deployment overhead. We also introduce TacAura, a dataset of synchronized vision, proprioception, and vision-based tactile signals across four representative contact-rich manipulation tasks. Experiments show that τ outperforms existing models and generalizes to unseen objects and scenes, delivering improved manipulation performance and robustness. Project Page: https://cocacola-lab.github.io/tau-Page/.