arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08638cs.ROcs.AI

CASD:用于多阶段机器人操作的分块对齐语义蒸馏

CASD: Chunk-Aligned Semantic Distillation for Multi-StageRobot Manipulation

Tinghe Ding, Jiahao Li, He Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出CASD方法,利用离线视觉-语言模型为动作分块生成加权语义目标,训练策略提升多阶段机器人操作成功率,在多个基准上超越参考方法。

中文摘要 AI 辅助

一个动作分块可以跨越操作任务的多个阶段,然而其第一步的标签仅描述当前阶段。我们提出了分块对齐语义蒸馏(CASD),该方法为整个动作分块推导语义目标。一个离线的视觉-语言模型将演示分割为已描述的阶段。这些阶段在每个动作分块内的占用率决定了一个加权语义目标,包括阶段之间的转换。一个CASD生成器学习从当前观察、机器人状态和任务指令中预测该目标。然后我们冻结生成器,并训练一个以其预测为条件的策略。语义分支在每次策略查询时运行一次,无需在线VLM调用或推理轨迹解码。在标注的LIBERO训练片段上,教师匹配对于单阶段和跨边界分块均高于随机水平。我们在四个基准上评估了三种Fast-WAM变体和一个DreamZero集成,包括LIBERO-Plus上的分布偏移。与已发表的参考相比,IDM+CASD在LIBERO上达到了98.9%的平均成功率,而参考为98.0%,而Uncond则低于其参考。Joint+CASD在RoboTwin 2.0上达到了93.0%,而参考为90.6%,DreamZero+CASD在MolmoSpaces四类别操作平均上达到了47.9%,而参考为40.7%。性能因骨干集成而异。

英文摘要

An action chunk can span several stages of a manipulation task, yet a label for its first step describes only the current stage. We introduce Chunk-Aligned Semantic Distillation (CASD), which derives semantic targets for entire action chunks. An offline vision--language model segments demonstrations into described stages. Their occupancy within each action chunk determines a weighted semantic target, including transitions between stages. A CASD generator learns to predict this target from the current observation, robot state, and task instruction. We then freeze the generator and train a policy conditioned on its predictions. The semantic branch runs once per policy query, without online VLM calls or reasoning-trace decoding. Teacher matching on annotated LIBERO training episodes is above chance for both single-stage and boundary-crossing chunks. We evaluate three Fast-WAM variants and a DreamZero integration across four benchmarks, including distribution shifts on LIBERO-Plus. Compared with published references, IDM+CASD reaches 98.9\% versus 98.0\% average success on LIBERO, while Uncond falls below its reference. Joint+CASD reaches 93.0\% versus 90.6\% on RoboTwin 2.0, and DreamZero+CASD reaches a 47.9\% four-category MolmoSpaces manipulation average versus 40.7\%. Performance varies across backbone integrations.

发表机构

  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

↑