arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

X2Real:面向真实世界通用策略的广泛仿真基准

X2Real: an eXtensive simulation benchmark for real-world generalist policies

Lian Ruan, Jade Yang, Sherphylan Gao, Felix Gao, Kyson Liang, Galen Liu, Ligo Wu, Lane Jin, Guu Gu, Bevan Xie, Cloud Yan, Zongzi Yuan, Kino Luo, Emma Chen, Shuwen Chen, Yang Ping, Miles Guo, Rain Sun, Kayden Zhang, Alex Du, Ruihai Wu, Liang Hao, Zhaoshuo Li, Roy Gan, Hao Wang, Qian Wang

arXiv 2609.27449首次发表:更新:

AI 中文总结

X2Real提出可进化的仿真基准,通过校准保真度、多样性与公平性,实现0.84的仿真-真实相关性,支持通用机器人操作策略的可靠评估与发展。

AI 中文摘要

通用机器人操作策略发展迅速,然而其可靠评估仍因现有仿真基准的根本缺陷而面临挑战:突出的仿真到现实差距、狭窄的任务覆盖范围,以及由模糊的训练-测试流程导致的不公平评估。先前的工作仅部分解决了这些问题,缺乏同时具备保真性、多样性和公平性,而静态基准设计无法支持长期策略发展。我们提出X2Real,一个基于Nvidia Isaac Lab-Arena的可进化仿真基准,用于忠实评估机器人操作策略的真实世界性能。遵循三个核心原则(保真性、多样性和公平性),X2Real校准仿真视觉和物理属性以与真实硬件对齐,实现了仿真与真实机器人评估结果之间0.84的线性相关性。它包含一个全面的分类体系,涵盖10个能力维度和44个分层长时程任务,覆盖基本操作技能和高级能力,如视觉定位、语言理解和双臂控制。我们进一步采用多轴域随机化和严格分离的训练-评估流程,以减轻基准利用并确保可信评估。借助自定义物理领域特定语言,Mana仿真生态系统支持模块化任务设计和迭代性能分析,以及近300小时的带注释仿真轨迹数据集。X2Real提供了一个保真、多样且公平的进化评估基础设施,有效弥合仿真到现实评估差距,并支持通用机器人操作策略的进步。

英文摘要

Generalist robot manipulation policies have developed rapidly, yet their reliable evaluation remains challenging due to fundamental flaws in existing simulation benchmarks: prominent sim-to-real gaps, narrow task coverage, and unfair evaluation caused by ambiguous training-test pipelines. Prior works only partially resolve these issues and lack simultaneous faithfulness, diversity, and fairness, while static benchmark designs fail to sustain long-term policy development. We present X2Real, an evolvable simulation benchmark for faithfully evaluating the real-world performance of robotic manipulation policies based on Nvidia Isaac Lab-Arena. Following three core principles (faithfulness, diversity, and fairness), X2Real calibrates simulation visual and physical properties to align with real hardware, achieving a 0.84 linear correlation between simulated and real-robot evaluation results. It features a comprehensive taxonomy with 10 capability dimensions and 44 hierarchical long-horizon tasks, covering basic manipulation skills and advanced capacities such as visual grounding, language understanding, and bimanual control. We further adopt multi-axis domain randomization and strictly disjoint training-evaluation pipelines to mitigate benchmark exploitation and ensure credible evaluation. Powered by a custom physical domain-specific language, the Mana simulation ecosystem supports modular task design and iterative performance analysis, alongside a nearly 300-hour annotated simulation trajectory dataset. X2Real offers a faithful, diverse, and fair evolving evaluation infrastructure, effectively bridging the sim-to-real evaluation gap and supporting the advancement of generalist robotic manipulation policies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑