arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GPUSimBench:迈向具身人工智能中可扩展且可靠的GPU加速模拟器

GPUSimBench: Towards Scalable and Reliable GPU-Accelerated Simulators in Embodied AI

Huzhenyu Zhang, Shenghai Yuan, Wenrui Yan, Li Ma, Hengjie Li, Jingcheng Pang, Dmitry Yudin

arXiv 2607.13059首次发表:更新:

发表机构

Shanghai AI Laboratory; MIRAI; Nanyang Technological University; Nanjing University(上海人工智能实验室; 未来人工智能研究所; 南洋理工大学; 南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究具身人工智能中GPU加速模拟器问题,通过GPUSimBench工具,建立物理基础评估、基准测试并行可扩展性,揭示并量化GPU批处理执行的不确定性,确定模拟器堆栈随机经验模式,强调无界扩展对可重复性的影响。

AI 中文摘要

数据驱动的具身人工智能正迅速转变为通过大规模并行模拟来扩展训练的范式,其中GPU加速模拟器是基础数据基础设施。然而,随着计算吞吐量的扩展,并行效率、物理保真度和执行确定性之间的潜在权衡仍未得到充分研究,阻碍了可靠机器人学习的发展。本文通过引入GPUSimBench揭示了主流基于GPU的机器人模拟器(如Isaac Lab、Genesis)的隐藏局限性,该工具专注于可扩展性、物理一致性和计算确定性。首先,GPUSimBench通过一个受控的斜面任务建立物理基础评估,量化模拟动力学与其现实世界对应物之间的分布对齐。其次,我们通过测量跨扩展环境数量的吞吐量和内存占用情况来基准测试并行可扩展性。至关重要的是,除了标准性能指标外,我们还揭示并量化了GPU批处理执行引入的固有不确定性,其特征是即使在相同初始条件下,运行到运行和环境间也存在显著差异。最后,我们确定了当前模拟器堆栈中的四种随机经验模式,强调无界扩展在没有明确约束时会损害可重复性。

英文摘要

Data-driven embodied AI is rapidly transitioning into a paradigm that scales training through massively parallel simulation, where GPU-accelerated simulators serve as the foundational data infrastructure. However, as computational throughput scales, the underlying trade-offs between parallel efficiency, physical fidelity, and execution determinism remain largely unexamined, hindering the development of reliable robot learning. In this paper, we expose the hidden limits of mainstream GPU-based robotic simulators (e.g., Isaac Lab, Genesis) by introducing GPUSimBench, which focuses on scalability, physical consistency, and computational determinism. First, GPUSimBench establishes a physical grounding evaluation with a controlled inclined-plane task, quantifying the distributional alignment between simulated dynamics and their real-world counterparts. Second, we benchmark parallel scalability by measuring throughput and memory footprints across scaling environment counts. Crucially, beyond standard performance metrics, we unveil and quantify the inherent non-determinism introduced by GPU-batched execution, characterized by significant run-to-run and inter-environment variability even under identical initial conditions. Finally, we identify four empirical regimes of stochasticity within current simulator stacks, highlighting that unbounded scaling can compromise reproducibility without explicit constraints.

CommentsAccepted by IROS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑