arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.06988cs.ROcs.AI

WAM-TTT:在测试时通过观察人类行为来引导世界行动模型

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Yusen Feng, Bingchen Han, Jiangran Lyu, Kai Liu, Yixin Zheng, Yuxuan Wan, Weiheng Liu, Sun Han, Ruiqin Li, Yulong Zhang, Fangfu Liu, Xuesong Shi, Libin Liu, Yiz… 展开作者

Yusen Feng, Bingchen Han, Jiangran Lyu, Kai Liu, Yixin Zheng, Yuxuan Wan, Weiheng Liu, Sun Han, Ruiqin Li, Yulong Zhang, Fangfu Liu, Xuesong Shi, Libin Liu, Yizhou Wang, Zhizheng Zhang, He Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究旨在解决引导机器人基础模型的难题,提出WAM-TTT框架,通过自监督视频预测吸收人类视频到自适应内存,经元训练阶段对齐人机行为,测试时仅用未标记人类视频,实现高效可重复引导且优于基线。

中文摘要 AI 辅助

引导机器人基础模型适应新任务变体或用户偏好行为具有挑战性,通常需要额外的机器人演示、特定任务微调或长上下文条件设定。我们提出了WAM-TTT,一个在测试时从原始人类视频引导世界行动模型的训练框架。它通过自监督视频预测将人类视频吸收到冻结的WAM内的轻量级自适应内存中,而非将其视为模仿轨迹。我们引入元训练阶段,使用配对的人机数据和键值内存重建目标使人类演示与机器人行为对齐。测试时仅需未标记的人类视频来调整内存,预训练的WAM保持冻结。这实现了无需机器人动作、人工标注或特定任务微调的高效且可重复使用的引导,同时保留基础模型的泛化能力。大量实验表明,WAM-TTT在各种操纵任务和泛化设置中始终优于上下文人类视频条件基线方法。

英文摘要

Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperforms in-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.

发表机构

  • Peking University(北京大学)
  • Galbot
  • CASIA(中国科学院自动化所)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑