发表机构
Texas A&M University(德克萨斯A&M大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ForeTac-VLA提出基于预测的触觉-视觉-语言融合模型,预测未来触觉状态以引导动作生成,在接触丰富操作任务中平均成功率95%,显著优于现有基线。
AI 中文摘要
视觉-语言-动作(VLA)模型在机器人操作中展现出强大的能力,然而其对视觉感知的依赖限制了在接触丰富环境中的鲁棒性,在这些环境中,关键的物理交互状态可能无法通过视觉观察到。现有的触觉增强VLA方法利用观测到的触觉反馈改善了物理接地,但大多数方法仍主要是反应式的,而非显式地建模接触如何演变。因此,我们提出了ForeTac-VLA,一种基于预测的触觉-视觉-语言融合模型,通过预测未来的触觉状态来引导动作生成。具体而言,ForeTac-VLA将最近的触觉观测编码为时间表示,并通过双向交叉注意力将其与视觉-语言特征集成。此外,一个基于Transformer的预测模块预测多步未来的触觉状态,使模型能够对观测到的和预期的接触进行联合推理。最后,融合的多模态表示和预测的未来触觉状态被输入到VLA骨干网络中,以条件化动作生成。为了稳定训练,当早期预测不可靠时,采用从真实值到预测的课程学习策略。在四个真实世界的接触丰富操作任务中,ForeTac-VLA实现了95%的平均成功率,比微调的VLA模型高出36.25个百分点,比最先进的触觉增强VLA基线高出超过22个百分点。ForeTac-VLA在低光照和视觉杂乱条件下也保持了强劲的性能。视频演示可在以下网址找到:此https URL
英文摘要
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, yet their reliance on visual perception limits robustness in contact-rich environments, where critical physical interaction states may not be visually observable. Existing tactile-enhanced VLA methods improve physical grounding using observed tactile feedback, but most remain largely reactive rather than explicitly modeling how contact may evolve. Therefore, we propose ForeTac-VLA, a forecasting-based tactile-vision-language fusion model that predicts future tactile states to guide action generation. Specifically, ForeTac-VLA encodes recent tactile observations into temporal representations and integrates them with vision-language features through bidirectional cross-attention. Further, a transformer-based forecasting module predicts multi-step future tactile states, enabling the model to reason jointly over observed and anticipated contact. Finally, the fused multimodal representations and predicted future tactile states are fed into the VLA backbone to condition action generation. To stabilize training, a ground-truth-to-prediction curriculum is employed when early forecasts are unreliable. Across four real-world contact-rich manipulation tasks, ForeTac-VLA achieves an average success rate of 95%, outperforming the fine-tuned VLA model by 36.25 percentage points and state-of-the-art tactile-enhanced VLA baselines by over 22 percentage points. ForeTac-VLA also maintains strong performance under low-illumination and visually cluttered conditions. Video demonstrations can be found on https://foretac-vla.github.io/
Comments8 pages, 7 figures