发表机构
AXXX; MIRIAI; HSE University(AXXX; MIRIAI; 高等经济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MIKASA-Robo-VLA基准,含90个语言操控任务,多数需记忆隐藏线索,发布22,500条轨迹,基线π0.5在14任务上成功率0.211,揭示VLA模型记忆局限。
AI 中文摘要
视觉-语言-动作策略通常仅观察一个或少数几个最近的帧,这使得评估它们如何在任务过程中利用消失的信息变得困难。我们引入了MIKASA-Robo-VLA,一个包含90个语言条件化操控任务的基准测试。除10个任务外,其余所有任务都隐藏了动作所依赖的线索。那10个任务是反应式控制。MIKASA-Robo是它所重建的套件,包含32个任务,并且仅在代表性的VLA子集中使用语言。在这里,每个任务都提供指令,而记忆依赖任务隐藏任务相关线索,反应式控制则保持线索可用。对于70个任务,环境阶段时间指定了信息间隔,其中28个任务的间隔超过了我们所调查的最宽固定上下文VLA的16帧窗口。该间隔仅计算线索被证明缺失的时段,而非策略必须保留线索的完整持续时间,因此每个记忆依赖任务在构造上仍然需要记忆,包括那些测量间隔较短的任务。我们在RLDS和LeRobotDataset v3中发布了跨越10种记忆类型的22,500条专家轨迹。一个参考的π0.5基线,使用当前图像和本体感觉,但没有观察历史或显式记忆模块,在14个任务上进行了微调,实现了0.211±0.044的平均任务成功率。它在评估的Long-split任务上的较低成功率受到开环分块和该子集中所代表的记忆类型的混淆。项目页面:此https URL
英文摘要
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference $π_{0.5}$ baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 $\pm$ 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/
Comments57 pages, 39 figures, 38 tables