arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TACO:作为可扩展VLA训练后自校正器的触觉世界模型

TACO: TActile World Model as a Self-COrrector for Scalable Robot Policy Post-Training

Shengbang Liu, Yueru Jia, Yuyang Yan, Jiaming Liu, Xinran Zhang, Qiuxuan Feng, Yandong Guo, Shiji Zhou, Boxin Shi, Shanghang Zhang

arXiv 2607.02840首次发表:更新:

发表机构

State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; AI 2 Robotics; Sun Yat-sen University; Beihang University(北京大学计算机科学学院多媒体信息处理技术国家重点实验室; 人工智能与机器人研究所; 中山大学; 北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究VLA模型在接触丰富任务中的问题,提出TACO框架,通过触觉感知世界模型驱动可扩展VLA训练后校正,结合多种方法提升策略,实验表明其能显著提高成功率。

AI 中文摘要

视觉-语言-动作(VLA)模型在机器人操作中泛化性良好,但在接触丰富任务中存在问题。触觉感知的训练后校正可改善恢复,但人工干预成本高。TACO是用于接触丰富操作中可扩展VLA训练后校正的触觉感知世界模型驱动框架,实验表明其能显著提升成功率。

英文摘要

Vision-Language-Action models and World Action Models have shown promising generalization in robotic manipulation but remain fragile in contact-rich tasks, where contact perturbations can cause failures that are difficult to detect from vision alone. Corrective post-training with tactile feedback can improve recovery, but scaling such supervision through human intervention is costly. World models can synthesize additional training data, yet vision-only generation may produce visually plausible but contact-inconsistent trajectories. We therefore introduce TACO, a scalable robot policy post-training framework built on a compositional tactile world model. Given real rollouts, TACO follows a Recognize--Imagine--Label loop: an inverse dynamics and value model identifies failure-adjacent states using progress estimates, a visuo-tactile generation model imagines local corrections by jointly generating video and tactile sequences, and the inverse dynamics and value model labels them with corrective actions and progress scores. Candidates are filtered for kinematic feasibility and tactile plausibility, then selected by predicted progress gain. TACO aggregates demonstrations, real rollouts, and selected corrections for iterative post-training. It combines knowledge-insulated tactile adaptation with CFG-RL using binary advantage labels while keeping the pretrained VLM backbone fixed. Experiments on real-world tasks show that TACO improves the average task score from 0.375 to 0.825 after two post-training iterations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑