发表机构
Nanyang Technological University; Agency for Science, Technology and Research (A*STAR)(南洋理工大学; 新加坡科技研究局)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MotionForge提出首个大规模仿真基准与数据生成流水线,含40个动态任务(17个长时程),支持单因素及联合域偏移评估与延迟感知执行,揭示现有策略在联合域偏移下的显著局限。
AI 中文摘要
近年来,基于学习的机器人策略取得了令人鼓舞的进展,然而这些策略主要是在静态或准静态环境中进行评估。在动态操作中,当机器人进行感知、推理和行动时,对象和场景会持续演变。然而,现有的动态仿真基准大多侧重于具有简单运动模式的短时程、反应式交互,并且对域偏移下的系统性评估以及与模型无关的实时执行协议支持有限。为了弥合这些差距,我们提出了MotionForge,这是第一个专门用于联合评估动态操作中域偏移和长时程交互的大规模仿真基准和数据生成流水线。MotionForge包含40个动态交互任务,涵盖11种不同的运动模式,并专门支持17个长时程任务。我们的基准引入了两个关键创新:(1)一个系统性评估协议,用于评估策略在单因素(例如,仅背景偏移)和联合域偏移(例如,对象、背景、光照和速度同时偏移)下的鲁棒性;(2)一个解耦的、延迟感知的执行协议,其中环境独立于策略推理时间持续演变。在我们基准上对代表性通用机器人策略的广泛评估揭示了它们在联合域偏移下的显著局限性。这些发现暴露了当前策略能力与在域偏移下对动态对象进行鲁棒长时程操作要求之间的关键差距,使MotionForge成为具身智能未来研究的综合测试平台。
英文摘要
Recent advances in learning-based robot policies have demonstrated promising progress, yet they are predom- inantly evaluated in static or quasi-static environments. In dynamic manipulation, objects and scenes continuously evolve while the robot perceives, reasons, and acts. However, recent dynamic simulation benchmarks largely focus on short-horizon, reactive interactions with simple motion patterns and offer limited support for both systematic evaluation under domain shifts and model-agnostic real-time execution protocols. To bridge these gaps, we introduce MotionForge, the first large- scale simulation benchmark and data-generation pipeline tailored to jointly evaluate domain shifts and long-horizon interaction in dynamic manipulation. MotionForge comprises 40 dynamic interaction tasks spanning 11 distinct motion patterns, with dedicated support for 17 long-horizon tasks. Our benchmark introduces two key novelties: (1) a systematic evaluation protocol for assessing policy robustness under both single-factor (e.g., only backgrounds shift) and joint domain shifts (e.g., simultaneous shifts of objects, backgrounds, lighting, and speed); and (2) a decoupled, latency-aware execution protocol where the environ- ment continuously evolves independently of policy inference time. Extensive evaluations of representative general-purpose robot policies on our benchmark reveal substantial limitations under joint domain shifts. These findings expose a critical gap between current policy capabilities and the requirements of robust long- horizon manipulation of dynamic objects under domain shifts, establishing MotionForge as a comprehensive testbed for future research in embodied AI.
Comments9 pages