arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WALA:从带动作标签的演示和无动作视频中学习可执行的潜在动作

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

Jiahao Liu, Yupeng Zheng, Zhongpu Xia, Shuai Tian, Huangrui Li, Yuhang Zheng, Ning Ma, ShangQing Zhou, Xiaotian Liu, Jing Li, Yixian Li, Haoran Li, Dongbin Zhao

arXiv 2607.11397首次发表:更新:

发表机构

CASIA; Anyverse Dynamics; UCAS; NUS; XJYLU(中国科学院自动化所; Anyverse动力学公司; 中国科学院大学; 新加坡国立大学; 西安交通大学利物浦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究从带动作标签演示和无动作视频学习可执行潜在动作的问题,提出WALA框架,先预训练语义几何潜在动作模型,策略训练时编码器和解码器协同,由多方面联合监督潜在动作,实验证明其提升了机器人策略性能和泛化能力。

AI 中文摘要

可泛化的机器人策略通常依赖带动作标签的机器人演示,但其收集成本高且难以扩展。大规模人类和机器人视频虽有丰富物理交互但缺乏可执行机器人动作标签。我们提出WALA框架,它能从带动作标签的演示和无动作视频中学习可执行的潜在动作。先通过对当前与稀疏采样的未来观测间的演化建模,从视频中预训练语义几何潜在动作模型,预测DINOv3特征空间和密集深度空间中的未来增量。在策略训练时,预训练编码器提供稳定潜在动作目标,解码器作为可训练的潜在世界模型。潜在动作由机器人动作预测等联合监督。实验表明WALA在RoboTwin上性能强劲,在RoboCasa上取得新的最先进结果,且提升了实际操作任务中的策略性能和泛化能力。

英文摘要

Human videos provide rich information about object manipulation at a scale difficult to reproduce with robots, but often lack action annotations that can directly supervise robot policies. Realizing this potential requires distinguishing interaction-relevant changes from static appearance and connecting the learned interaction knowledge to executable robot actions. We present WALA, a framework named for World- and Action-supervised Latent Actions. Its semantic-geometric latent action model (LAM) learns latent actions by encoding and predicting semantic and geometric changes, emphasizing interactions over static appearance while retaining spatial detail. During robot policy training, the LAM encoder and decoder provide guidance through distillation and visual prediction, respectively. Together with robot action supervision, these signals shape policy-generated latent actions to support world prediction and capture the information needed to directly generate executable robot actions. Because LAM supervision requires no action labels, WALA enables co-training on action-labeled robot demonstrations and task-relevant action-free videos. In simulation experiments, our method achieves 92.33% average success on RoboTwin and 75.2% on RoboCasa-GR1-Tabletop, exceeding the previous state of the art on RoboCasa by 5.0 percentage points. WALA also shows strong real-robot performance, further enhanced by task-relevant action-free human videos that improve data efficiency and support target-task zero-shot and few-shot transfer.

CommentsProject page: https://liujiahao2077.github.io/WALA.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑