LIBERO-MAX:当世界变化时,机器人策略会适应吗?
LIBERO-MAX: Do Robot Policies Adapt When the World Changes?
- Tulane University(杜兰大学)
- New York University(纽约大学)
- UIUC(伊利诺伊大学厄巴纳-香槟分校)
- Stanford University(斯坦福大学)
- CMU(卡内基梅隆大学)
- University of Rochester(罗切斯特大学)
- Nanyang Technological University(南洋理工大学)
- MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
LIBERO-MAX基准通过8,000个配对案例评估机器人策略在任务中途变化下的适应性,发现事件导致成功率下降11.0-25.7个百分点,揭示了共同脆弱性,为诊断和进展测量提供可复现测试平台。
AI中文摘要:
机器人通常必须在目标移动、视角改变或出现障碍物后继续执行任务,尽管它们之前的观察和已承诺的动作反映了之前的场景。许多模拟鲁棒性基准在重置时固定外部条件,使得这一时间上的挑战未被充分研究。我们引入了LIBERO-MAX,一个包含8,000个配对案例的基准,涵盖几何、观察、外观、杂乱和路径的八种变化类型。每对比较在有和没有任务中途事件的情况下的任务执行,保持任务、初始状态、策略种子和事件前动作序列固定。这种受控比较区分了与事件相关的性能下降与在没有变化时已经存在的失败。在十四个当前的VLA、混合和世界动作策略中,事件使成功率降低了11.0-25.7个百分点。事件概况揭示了在几何和观察变化上的共同脆弱性,而策略家族排名相互交织。相机控制表明,鲁棒性既反映了在变化条件下的能力,也反映了遇到这些变化的轨迹;改变查询频率并不能消除差距。总之,配对协议和时间诊断使LIBERO-MAX成为一个可复现的测试平台,用于诊断在执行中变化下的失败,并衡量在变化世界中保持有效的机器人策略的进展。
英文摘要:
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.