MINERVA:操控策略需要多小才能解决LIBERO任务?
MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?
查看机构详情
- Graduate School of Engineering, The University of Tokyo(东京大学工学研究科)
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对LIBERO操作基准所需的模型容量尚不明确的问题,本研究提出MINERVA系列紧凑视觉运动策略以测算任务特定容量下限,发现0.54M参数策略在LIBERO任务上表现优异且效率极高,为部署高效机器人策略提供了容量感知设计的依据。
中文摘要 AI 辅助
拥有数十亿参数的视觉-语言-动作(VLA)模型目前在LIBERO操作基准中占据主导地位,但该基准实际所需的模型容量仍不明确。我们提出MINERVA(MINimal Efficient Robotic Vision-Action policy,即最小高效机器人视觉-动作策略),这是一组刻意设计的紧凑视觉运动策略,用于测算该基准的任务特定容量下限。一款参数规模为0.54M的策略在四个标准LIBERO套件的2000次rollout中实现了95.1%的平均成功率,尽管参数数量比已报道的LeRobot π₀.₅少7700倍,但仅比后者低2.4个百分点。性能在参数接近1M时趋于饱和,低于0.25M时则会崩溃。在广泛的架构、训练和推理扫描中,仅动作块长度和视觉容量始终超出±1个百分点的训练种子区间。在三个种子下,Flow匹配与直接L1回归相比未表现出可检测的优势,而回归在GPU上的速度最高可达3.8倍。任务ID置换探针显示,标准LIBERO指令条件主要在记忆的任务中进行选择:仅改变任务ID映射就会使成功率降至接近随机水平。采用相同方案在89项LIBERO-90任务中实现了94.6%的成功率,而LIBERO-Plus扰动将性能降至46-56%,对光度偏移几乎没有鲁棒性。这款0.54M参数的策略在笔记本电脑CPU上每控制步骤的重规划时间为每块5-9毫秒,比SmolVLA快113倍,比π₀.₅快1400倍,且无需GPU。这些结果首次对LIBERO的任务特定容量下限进行了实证估计,并为部署高效机器人策略的容量感知设计和蒸馏提供了动机。
英文摘要
Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, but the model capacity actually required by the benchmark remains unclear. We introduce MINERVA (MINimal Efficient Robotic Vision-Action policy), a family of deliberately compact visuomotor policies designed to measure this task-specific capacity floor. A 0.54M-parameter policy achieves 95.1% average success over 2,000 rollouts on the four standard LIBERO suites, only 2.4 points below the reported LeRobot $π_{0.5}$ result despite using 7,700$\times$ fewer parameters. Performance saturates near 1M parameters and collapses below 0.25M. Across broad architectural, training, and inference sweeps, only action-chunk length and vision capacity consistently exceed a $\pm$1-point training-seed band. Flow matching provides no detectable advantage over direct L1 regression across three seeds, while regression is up to 3.8$\times$ faster on GPU. A task-ID permutation probe shows that standard LIBERO instruction conditioning primarily selects among memorized tasks: changing only the task-ID mapping reduces success to near chance. The same recipe achieves 94.6% success across 89 LIBERO-90 tasks, while LIBERO-Plus perturbations reduce performance to 46--56%, with near-zero robustness to photometric shifts. The 0.54M policy replans every control step in 5--9 ms per chunk on a laptop CPU, 113$\times$ faster than SmolVLA and 1,400$\times$ faster than $π_{0.5}$, without a GPU. These results establish a first empirical estimate of LIBERO's task-specific capacity floor and motivate capacity-aware design and distillation for deployment-efficient robot policies.