发表机构
University of Southern California; Mitsubishi Electric Research Laboratories (MERL); Mitsubishi Electric(南加州大学; 三菱电机研究实验室(MERL); 三菱电机)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TacSushi通过触觉门控融合与仅训练的未来后果监督,提升灵巧寿司操作在分布内外任务中的成功率。
AI 中文摘要
灵巧的食物操作需要在变形、遮挡和不确定接触条件下进行控制。我们提出了TacSushi,一种基于触觉接地、基于Cosmos3的世界-动作策略,它在基于当前观测行动的同时,从记录的未来后果中学习。骨干网络编码当前的RGB、语言和手部状态,特征级门控融合将指尖触觉特征纳入动作表示。在训练期间,以示范动作块为条件的解码器预测记录的未来视觉观测、任务进度、相对接触风险和触觉摘要;该解码器在部署时被移除。失败的试验提供后果监督,但其动作被排除在模仿之外。我们在340次成功和50次失败的实机试验上训练TacSushi,并在三个分布内任务和两个分布外食材变体的600次独立 rollout 中比较六种方法。为了评估超出单一几何阈值的食品质量,我们使用锚定视觉质量协议对终端结果进行评分,该协议对每次 rollout 的五个人类评分和三个视觉语言模型评分进行等权平均。完整的TacSushi实现了68.3%的平均分布内成功率和37.5%的分布外成功率,相比之下,没有未来后果监督的为36.7%/10.0%,而直接触觉拼接代替门控融合的为25.0%/17.5%。这些比较支持了特征级门控触觉融合和仅训练预测监督的互补益处。
英文摘要
Dexterous food manipulation requires control under deformation, occlusion, and uncertain contact. We present TacSushi, a tactile-grounded, Cosmos3-based world-action policy that learns from recorded future consequences while acting on current observations. The backbone encodes current RGB, language, and hand state, and feature-wise gated fusion incorporates fingertip tactile features into the action representation. During training, a decoder conditioned on demonstrated action chunks predicts logged future visual observations, task progress, relative contact risk, and tactile summaries; this decoder is removed at deployment. Failed trials provide consequence supervision, but their actions are excluded from imitation. We train TacSushi on 340 successful and 50 failed real-robot trials and compare six methods in 600 separate rollouts across three in-distribution tasks and two out-of-distribution ingredient variants. To assess food quality beyond a single geometric threshold, we score terminal outcomes using an anchored visual-quality protocol that equally weights five human ratings and three vision-language-model ratings per rollout. Full TacSushi achieves 68.3% average in-distribution success and 37.5% out-of-distribution success, compared with 36.7%/10.0% without future-consequence supervision and 25.0%/17.5% with direct tactile concatenation in place of gated fusion. These comparisons support complementary benefits of feature-wise gated tactile fusion and training-only predictive supervision.
Comments8 pages, 5 figures