arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模仿者游戏:超越动作预测的机器人模仿能力基准测试

The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

Xunzhe Zhou, Yiyang Cai, Fengyi Wang, Ran Ju, Hanxiang Ren, Ruizhe Liu, Yu Zhang, Qian Luo, Feng Chen, Pei Zhou, Yi Ma, Yanchao Yang

arXiv 2608.22301首次发表:更新:

发表机构

The University of Hong Kong; TranscEngram; Fudan University; Zhejiang University(香港大学; TranscEngram; 复旦大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出四级基准《模仿者游戏》及配套数据集、评估平台,发现功能替代是机器人意图层面模仿的关键障碍,微调IG-10K预训练模型可显著提升性能。

AI 中文摘要

人类在意图层面进行模仿:给定演示后,我们会推断其目标,并利用手头的任何工具、物体和布局来完成该目标。相比之下,当前的机器人策略从视觉输入和语言指令中学习观测到动作的映射,未明确推断演示的任务。因此,从人类视频中学习在很大程度上仍停留在轨迹层面:模型能在几乎相同的场景中复现动作,但仍难以模仿演示者的意图,而非仅仅是演示者的动作。我们推出《模仿者游戏》,这是一个四级基准(L0-L3),逐步拉大人类演示与机器人自身场景之间的差距,明确轨迹复现不再足够、需要任务理解的节点。我们将其与IG-10K配对,这是迄今为止最大的环境对齐的人机配对数据集,也是唯一在真实和模拟环境中覆盖全部四个级别的数据集(超过20000个配对片段、50多个任务、6个领域),以及Imitator Arena,一个用于盲法A/B人类评估的开放平台。在九个最先进的模型中,性能从L0到L2保持稳定,但在L3级崩溃,这表明功能替代——通过不同的物体 affordance 实现相同意图——是意图层面模仿的决定性障碍。以人类视频为条件的模型优于以字幕为条件的模型,但所有模型在未见过的任务上的零样本成功率均低于13%;仅用10个人机配对演示微调IG-10K预训练模型,即可获得大幅提升,且提升幅度随预训练规模增大而增加。项目网站和Imitator Arena的访问链接见this https URL。

英文摘要

Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at hand. Current robot policies instead learn observation-to-action mappings from visual inputs and language instructions, without explicitly inferring the demonstrated task. Learning from human video thus remains largely trajectory-level: models can replay motions in near-identical scenes, but still struggle to imitate what the demonstrator intends rather than merely what they do. We introduce The Imitator Game, a four-level benchmark (L0-L3) that progressively widens the gap between the human demonstration and the robot's own scene, isolating where trajectory replay ceases to suffice and task understanding becomes necessary. We pair it with IG-10K, the largest environment-aligned paired human-robot dataset to date and the only one instantiated across all four levels in both real and simulated settings (20,000+ paired episodes, 50+ tasks, 6 domains), and Imitator Arena, an open platform for blind A/B human evaluation. Across nine state-of-the-art models, performance is stable from L0 to L2 but collapses at L3, identifying functional substitution - achieving the same intent through a different object affordance - as the decisive barrier to intent-level imitation. Human-video-conditioned models outperform caption-conditioned ones, yet every model falls below 13% zero-shot success on unseen tasks; fine-tuning IG-10K-pretrained models with only $10$ paired human-robot demonstrations yields large gains that grow with pretraining scale. The project website and access to Imitator Arena are available at https://imitator-game.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑