LIBERO-VPro:机器人基础模型闭环视觉鲁棒性基准
LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models
浏览论文内容
中文总结 AI 辅助
提出LIBERO-VPro基准,系统评估机器人基础模型在视觉扰动下的闭环鲁棒性,揭示名义性能掩盖的视觉基础弱点,并表明鲁棒性多维且依赖架构。
中文摘要 AI 辅助
机器人基础模型在标准操作基准上取得了令人印象深刻的性能,然而这些评估通常假设在执行过程中视觉观测是干净、及时且一致的。我们提出了LIBERO-VPro,一个通过扰动执行过程中可用的视觉证据来系统评估机器人基础模型闭环视觉鲁棒性的基准。LIBERO-VPro涵盖四个互补维度,包括视觉证据退化、相机陈旧性、视觉源一致性和任务相关场景变化,跨越12个挑战类别、96个实验设置和3,296个任务条件案例。我们评估了三个视觉-语言-动作模型和三个世界-动作模型,涉及约196,000个模拟回合,并辅以在Franka Research 3上进行的200次真实世界部署。我们的结果表明,强大的名义性能可能掩盖视觉基础与适应方面的重大弱点。模型通常在严重的物体级遮挡下仍能保持成功,但当局部交互线索被破坏或熟悉的空间先验被违反时,性能急剧下降。它们对陈旧或缺失的观测也高度敏感,并且在改变的任务前提条件需要行为适应时表现挣扎。最后,VLA和WAM展现出不同的鲁棒性特征,表明视觉鲁棒性是多维且依赖架构的。LIBERO-VPro为开发能够在挑战性视觉条件下更可靠地基础化并调整其动作的机器人基础模型提供了系统的诊断框架。
英文摘要
Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic foundation models by perturbing the visual evidence available during execution. LIBERO-VPro covers four complementary dimensions, including Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task-Relevant Scene Variation, spanning 12 challenge categories, 96 experimental settings, and 3,296 task-condition cases. We evaluate three vision-language-action models and three world-action models over approximately 196,000 simulated episodes, complemented by 200 real-world rollouts on a Franka Research 3. Our results reveal that strong nominal performance can mask substantial weaknesses in visual grounding and adaptation. Models often remain successful despite severe object-level occlusion, yet degrade sharply when local interaction cues are disrupted or familiar spatial priors are violated. They are also highly sensitive to stale or missing observations and struggle when changed task preconditions require behavioral adaptation. Finally, VLAs and WAMs exhibit distinct robustness profiles, showing that visual robustness is multi-dimensional and architecture-dependent. LIBERO-VPro provides a systematic diagnostic framework for developing robotic foundation models that can more reliably ground and adapt their actions under challenging visual conditions.
发表机构
- Singapore Management University(新加坡管理大学)
- Princeton University(普林斯顿大学)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。