IndustrialVLA-Bench:开放机器人策略模型的可追踪多轴评估
IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models
- Griffith University(格里菲斯大学)
- Southern University of Science and Technology(南方科技大学)
- Jiangsu University(江苏大学)
- University of Southern Queensland(南昆士兰大学)
- The University of Sydney(悉尼大学)
- Norve Labs Inc.(Norve实验室公司)
- ACE Robotics(ACE机器人公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
IndustrialVLA-Bench在统一协议下评估六个VLA和WAM系统,揭示能力、鲁棒性及语言敏感性的差异,提供可追踪的对比证据。
AI中文摘要:
开放机器人策略日益遵循两种范式:视觉-语言-动作模型(VLAs)直接将观察和指令映射到动作,而世界-动作模型(WAMs)则将学习到的视频或世界动态纳入策略学习或动作生成中。尽管两者都针对相同的操作任务并代表替代性的设计选择,但它们通常在不同的评估协议下进行报告,导致其能力、鲁棒性、语言敏感性和部署成本之间的权衡不明确。我们提出了IndustrialVLA-Bench,一个在统一报告模式下的证据感知评估,涵盖六个已发布的VLA和WAM系统。它分别评估了在LIBERO上的清洁能力、在LIBERO-Plus上的非语言鲁棒性、在LIBERO-Para上的指令敏感性以及观察到的执行成本。报告的任务得分汇总了在固定检查点和推理配置下使用不同随机种子的三次完整评估。在所有六个系统中,清洁LIBERO平均值仅相差1.58分,而鲁棒性和释义摘要分别跨越14.62和31.08分。将每次比较限制在三个符合协议的系统上保留了效果(1.36、14.62和23.10分),因此这里报告的诊断分离不依赖于较弱的证据层级。我们还报告了每个系统的观察推理延迟、峰值内存、运行时模式和证据状态。符合协议、接近复现和待验证的条目保持明显分离;只有符合协议的条目支持严格比较。IndustrialVLA-Bench并不声称任一范式的普遍优越性,而是为在共享实际标准上比较发布的机器人策略提供可追踪的证据。代码和评估记录可在https://this https URL获取。
英文摘要:
Open robot policies increasingly follow two paradigms: vision-language-action models (VLAs) directly map observations and instructions to actions, whereas world-action models (WAMs) incorporate learned video or world dynamics into policy learning or action generation. Although both target the same manipulation tasks and represent alternative design choices, they are commonly reported under different evaluation protocols, leaving their capability, robustness, language sensitivity, and deployment-cost trade-offs unclear. We present IndustrialVLA-Bench, an evidence-aware evaluation of six released VLA and WAM systems under a unified reporting schema. It separately evaluates clean capability on LIBERO, non-language robustness on LIBERO-Plus, instruction sensitivity on LIBERO-Para, and observed execution cost. Reported task scores aggregate three complete evaluations with distinct random seeds under a fixed checkpoint and inference configuration. Across all six systems, clean LIBERO averages differ by only 1.58 points, whereas robustness and paraphrase summaries span 14.62 and 31.08 points. Restricting every comparison to the three protocol-faithful systems preserves the effect (1.36, 14.62 and 23.10 points), so the diagnostic separation reported here does not depend on the weaker evidence tiers. We additionally report observed inference latency, peak memory, runtime mode, and an evidence status for every system. Protocol-faithful, near-reproduction, and pending-verification entries remain visibly separated; only protocol-faithful entries support strict comparisons. Rather than claiming universal superiority of either paradigm, IndustrialVLA-Bench provides traceable evidence for comparing released robot policies on shared practical criteria. Code and evaluation records are available at https://github.com/xiaoqi-7/IndustrialVLA-Bench.