AI 中文总结
X2Real提出可进化的仿真基准,通过校准保真度、多样性与公平性,实现0.84的仿真-真实相关性,支持通用机器人操作策略的可靠评估与发展。
AI 中文摘要
通用机器人操作策略发展迅速,然而其可靠评估仍因现有仿真基准的根本缺陷而面临挑战:突出的仿真到现实差距、狭窄的任务覆盖范围,以及由模糊的训练-测试流程导致的不公平评估。先前的工作仅部分解决了这些问题,缺乏同时具备保真性、多样性和公平性,而静态基准设计无法支持长期策略发展。我们提出X2Real,一个基于Nvidia Isaac Lab-Arena的可进化仿真基准,用于忠实评估机器人操作策略的真实世界性能。遵循三个核心原则(保真性、多样性和公平性),X2Real校准仿真视觉和物理属性以与真实硬件对齐,实现了仿真与真实机器人评估结果之间0.84的线性相关性。它包含一个全面的分类体系,涵盖10个能力维度和44个分层长时程任务,覆盖基本操作技能和高级能力,如视觉定位、语言理解和双臂控制。我们进一步采用多轴域随机化和严格分离的训练-评估流程,以减轻基准利用并确保可信评估。借助自定义物理领域特定语言,Mana仿真生态系统支持模块化任务设计和迭代性能分析,以及近300小时的带注释仿真轨迹数据集。X2Real提供了一个保真、多样且公平的进化评估基础设施,有效弥合仿真到现实评估差距,并支持通用机器人操作策略的进步。
英文摘要
Generalist robot manipulation policies have developed rapidly, yet their reliable evaluation remains challenging due to fundamental flaws in existing simulation benchmarks: prominent sim-to-real gaps, narrow task coverage, and unfair evaluation caused by ambiguous training-test pipelines. Prior works only partially resolve these issues and lack simultaneous faithfulness, diversity, and fairness, while static benchmark designs fail to sustain long-term policy development. We present X2Real, an evolvable simulation benchmark for faithfully evaluating the real-world performance of robotic manipulation policies based on Nvidia Isaac Lab-Arena. Following three core principles (faithfulness, diversity, and fairness), X2Real calibrates simulation visual and physical properties to align with real hardware, achieving a 0.84 linear correlation between simulated and real-robot evaluation results. It features a comprehensive taxonomy with 10 capability dimensions and 44 hierarchical long-horizon tasks, covering basic manipulation skills and advanced capacities such as visual grounding, language understanding, and bimanual control. We further adopt multi-axis domain randomization and strictly disjoint training-evaluation pipelines to mitigate benchmark exploitation and ensure credible evaluation. Powered by a custom physical domain-specific language, the Mana simulation ecosystem supports modular task design and iterative performance analysis, alongside a nearly 300-hour annotated simulation trajectory dataset. X2Real offers a faithful, diverse, and fair evolving evaluation infrastructure, effectively bridging the sim-to-real evaluation gap and supporting the advancement of generalist robotic manipulation policies.